Concept · Chapter 16: Evaluation, Reliability, Safety & Interpretability
Sparse Autoencoders and Dictionary Learning
A sparse autoencoder learns to rewrite a model's activation as a sparse sum of many learned directions (a dictionary), so that each direction, unlike a neuron, tends to stand for one interpretable feature.
The problem
Because of superposition, a model's neurons mix many features, so neither neurons nor a handful of probe directions give a full, readable account of what a layer represents.
The solution
Train a wide autoencoder on the model's activations with a penalty that keeps only a few of its hidden units active per input; each hidden unit's decoder vector becomes a candidate feature direction, labelled by the inputs that activate it.
The consequence
Dictionary learning has found millions of interpretable features in production models and lets researchers turn them up or down to test what they do, though dictionaries are incomplete and their features need validating.
You should understand first
- The Turing Test
- Symbolic AI
- Logic and Rules
- Expert Systems
- Knowledge Representation
- The Knowledge-Acquisition Bottleneck
- From Rules to Learning
- Supervised, Unsupervised and Self-Supervised Learning
- Vectors
- Features, Labels and Tasks
- Loss Functions
- Derivatives and Gradients
- Gradient Descent
- Linear Regression
- Probability and Distributions
- Entropy
- Softmax
- Cross-Entropy Loss
- Logistic Regression
- Hand-Crafted Features vs Learned Features
- Dot Product
- The Perceptron
- Activation Functions
- The Artificial Neuron
- Matrix Multiplication
- Multilayer Perceptron (MLP)
- Representation Learning
- Embeddings
- Probing Representations
- Superposition and Polysemantic Neurons
- Expected Value and Variance
- Sampling and Uncertainty
- Generalization, Overfitting and Underfitting
- Regularization
- Sparse Autoencoders and Dictionary Learning
Undoing superposition
If a layer's activation is a sum of a few active features out of many, then finding those features is a dictionary learning problem: learn a large set of directions (with much bigger than the layer width) such that every activation is approximately a sparse combination of them.
A sparse autoencoder does this with a neural network:
The L1 penalty pushes most to zero for any given input, so each input is explained by a few features. To label a feature, look at the inputs that activate it most.
From a small model to a production one
Towards Monosemanticity trained sparse autoencoders on 8 billion activations of the 512-neuron MLP layer of a one-layer Transformer, with dictionaries from 512 to 131,072 features, and studied a 4,096-feature run in detail; features such as Arabic script, DNA sequences, base64 and Hebrew were much more interpretable than the neurons. Established Cunningham and colleagues independently found sparse-autoencoder features more interpretable than other decompositions by automated measures. Established Scaling Monosemanticity trained dictionaries of about 1M, 4M and 34M features on the middle-layer residual stream of Claude 3 Sonnet, finding abstract, multilingual and multimodal features, including safety-relevant ones such as features related to deception and bias. EstablishedTesting that features are used
A feature that fires on Golden Gate Bridge text could be a correlate the model ignores. Clamping that feature to 10 times its maximum activation during the forward pass made Claude 3 Sonnet start to identify itself as the Golden Gate Bridge. Established Intervening and watching behaviour change is the same logic as causal tests of probes and circuits.
Limits
Dictionaries do not reconstruct activations perfectly, so some of the model's computation lives in the error term; many concepts have no clean feature at a given dictionary size; features split into finer ones as the dictionary grows; and their labels are human interpretations of activation examples. Whether sparse autoencoders find the model's "true" features is an open question. Active researchTiny example
Take the pentagon from the superposition lab: two hidden numbers, five feature directions. Given a hidden vector that is exactly feature 3's direction, any two directions could reconstruct it, but the sparsest explanation is "feature 3 alone". That preference for the sparsest explanation, at scale, is what the L1 penalty encodes.
Mini experiment
Open the interactive feature browser linked from Towards Monosemanticity and pick three features at random. For each, write your own label from the top activating examples before reading the authors'. How often do you agree, and what would you need to see to trust a label?
What to remember
- SAE: activation x → wide sparse code f → reconstruction x̂ = b + Σ fᵢ dᵢ. Loss = reconstruction error + λ·Σ|fᵢ|.
- Each decoder vector dᵢ is a feature direction; read it by the inputs that activate it most.
- Towards Monosemanticity (2023): a 512-neuron layer, dictionaries of 512 to 131,072 features; features for Arabic script, DNA, base64…
- Scaling Monosemanticity (2024): ~1M, 4M and 34M features on Claude 3 Sonnet's middle layer.
- Clamping the Golden Gate Bridge feature made the model talk as if it were the bridge: evidence features are used.
- Limits: reconstruction is imperfect, many concepts lack a clean feature, and labels are human interpretations.
Key papers
Sparse Autoencoders Find Highly Interpretable Features in Language Models
Hoagy Cunningham, Aidan Ewart et al. · 2023
Independent evidence, alongside Anthropic's, that sparse autoencoders can undo superposition in real language models.
Towards Monosemanticity: Decomposing Language Models With Dictionary Learning
Trenton Bricken, Adly Templeton et al. · 2023 · Transformer Circuits Thread
Showed that a sparse autoencoder can split a small model's polysemantic neurons into thousands of features that each mean one thing.
How to read it: Start with the summary and one detailed feature (the Arabic-script feature); the interactive feature browser linked from the article is the best way to get a feel for it.
Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet
Adly Templeton, Tom Conerly et al. · 2024 · Transformer Circuits Thread
Took dictionary learning from a toy-sized model to a production model and showed that clamping a feature changes behaviour.
How to read it: The 'Influence on Behavior' section is the evidence that features are used, not just correlated; the 'Feature Completeness' section is the honest limit (many concepts have no clean feature yet).