Henry Hu
About
Work
Blog
Links
Reading
Interesting reads
Things I'm reading and things I've finished.
Finding Alignments Between InterpretableCausal Variables and Distributed Neural Representations
arxiv.org
Read
Genealogy
— 4% read
Programming Refusal with Conditional Activation Steering
arxiv.org
Read
Genealogy
— 4% read
Australian Evidence-Based Clinical Practice Guideline for ADHD: Consumer Companion
adhdguideline.aadpa.com.au
Read
Genealogy
— 15% read
Tractatus Logico-Philosophicus (English)
books.wittgensteinproject.org
Read
Genealogy
— 8% read
Multi-Turn Credit Assignment with LLM Agents
hlfshell.ai
Read
Genealogy
— finished Aug 11, 2026
Instruction Tuning for Large Language Models: A Survey
arxiv.org
Read
Genealogy
— 4% read
Training Overview | RLHF and Post-Training Book by Nathan Lambert
rlhfbook.com
Read
Genealogy
— 3% read
My Research Process: Key Mindsets - Truth-Seeking, Prioritisation, Moving Fast — AI Alignment Forum
alignmentforum.org
Read
Genealogy
— 67% read
How I Think About My Research Process: Explore, Understand, Distill — AI Alignment Forum
alignmentforum.org
Read
Genealogy
— finished Jul 29, 2026
Tips for Empirical Alignment Research — LessWrong
lesswrong.com
Read
Genealogy
— finished Jul 29, 2026
Evolution of Concepts in Language Model Pre-Training
arxiv.org
Read
Genealogy
— 1% read
Your Language Model Secretly Contains Personality Subnetworks
arxiv.org
Read
Genealogy
— 1% read
Localizing Persona Representations in LLMs
arxiv.org
Read
Genealogy
— 14% read
MIB: A Mechanistic Interpretability Benchmark
arxiv.org
Read
Genealogy
— 8% read
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models
arxiv.org
Read
Genealogy
— 4% read
Emergent World Representations: Othello-GPT (Li et al., 2023)
arxiv.org
Read
Genealogy
— finished Jul 13, 2026
Actually, Othello-GPT Has A Linear Emergent World Representation — Neel Nanda
neelnanda.io
Read
Genealogy
— finished Jul 13, 2026
The Linear Representation Hypothesis and the Geometry of Large Language Models (Park, Choe, Veitch, 2024)
arxiv.org
Read
Genealogy
— finished Jul 13, 2026
Refusal in LLMs is mediated by a single direction — LessWrong
lesswrong.com
Read
Genealogy
— finished Jul 13, 2026
Steering Llama 2 via Contrastive Activation Addition (Panickssery et al., 2023)
arxiv.org
Read
Genealogy
— finished Jul 13, 2026
The Geometry of Truth: Emergent Linear Structure in LLM Representations (Marks & Tegmark, 2023)
arxiv.org
Read
Genealogy
— finished Jul 13, 2026
Linguistic Regularities in Continuous Space Word Representations
aclanthology.org
Read
Genealogy
— finished Jul 13, 2026
Steering GPT-2-XL by adding an activation vector (Turner et al., LessWrong 2023)
lesswrong.com
Read
Genealogy
— finished Jul 12, 2026
Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet (Templeton et al., Anthropic 2024)
transformer-circuits.pub
Read
Genealogy
— finished Jul 12, 2026
Zoom In: An Introduction to Circuits (Olah et al., Distill 2020)
distill.pub
Read
Genealogy
— finished Jul 12, 2026
Toy Models of Superposition (Elhage et al., Anthropic 2022)
transformer-circuits.pub
Read
Genealogy
— finished Jul 12, 2026
Towards Monosemanticity: Decomposing Language Models With Dictionary Learning (Bricken et al., Anthropic 2023)
transformer-circuits.pub
Read
Genealogy
— finished Jul 8, 2026
The Brothers Karamazov
gutenberg.org
Read
Genealogy
— finished Jul 8, 2026
The Vulnerable World Hypothesis (Nick Bostrom, 2019)
nickbostrom.com
Read
Genealogy
— finished Jul 6, 2026
International AI Safety Report 2026
internationalaisafetyreport.org
Read
Genealogy
— finished Jul 6, 2026
Instrumental convergence (Eliezer Yudkowsky, 2025)
arbital.com
Read
Genealogy
— finished Jul 6, 2026
Frontier Models are Capable of In-Context Scheming
apolloresearch.ai
Read
Genealogy
— finished Jul 6, 2026
Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs
arxiv.org
Read
Genealogy
— finished Jul 6, 2026
ML systems will have weird failure modes
bounded-regret.ghost.io
Read
Genealogy
— finished Jul 6, 2026
The OTHER AI Alignment Problem: Mesa-Optimizers and Inner Alignment
youtube.com
Watch @ 0:00
Genealogy
— finished Jul 6, 2026
Alignment faking in large language models
arxiv.org
Read
Genealogy
— finished Jul 6, 2026
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
arxiv.org
Read
Genealogy
— finished Jul 6, 2026
Existentialism
plato.stanford.edu
Read
Genealogy
— finished Jul 2, 2026
The World Inside Neural Networks
goodfire.ai
Read
Genealogy
— finished Jul 2, 2026