Skip to content
Road to Intelligence

Concept · Chapter 16: Evaluation, Reliability, Safety & Interpretability

Sparse Autoencoders and Dictionary Learning

FrontierUnderstand11 minDifficulty

A sparse autoencoder learns to rewrite a model's activation as a sparse sum of many learned directions (a dictionary), so that each direction, unlike a neuron, tends to stand for one interpretable feature.

The problem

Because of superposition, a model's neurons mix many features, so neither neurons nor a handful of probe directions give a full, readable account of what a layer represents.

The solution

Train a wide autoencoder on the model's activations with a penalty that keeps only a few of its hidden units active per input; each hidden unit's decoder vector becomes a candidate feature direction, labelled by the inputs that activate it.

The consequence

Dictionary learning has found millions of interpretable features in production models and lets researchers turn them up or down to test what they do, though dictionaries are incomplete and their features need validating.

Undoing superposition

If a layer's activation xx is a sum of a few active features out of many, then finding those features is a dictionary learning problem: learn a large set of directions d1,…,dFd_1, \dots, d_F (with FF much bigger than the layer width) such that every activation is approximately a sparse combination of them.

A sparse autoencoder does this with a neural network:

f=ReLU(Wencx+benc),x^=bdec+∑ifi di,L=∥x−x^∥2+λ∑i∣fi∣.f = \mathrm{ReLU}(W_{\text{enc}} x + b_{\text{enc}}), \qquad \hat x = b_{\text{dec}} + \sum_i f_i\, d_i, \qquad \mathcal{L} = \lVert x - \hat x\rVert^2 + \lambda \sum_i |f_i| .

The L1 penalty pushes most fif_i to zero for any given input, so each input is explained by a few features. To label a feature, look at the inputs that activate it most.

From a small model to a production one

Towards Monosemanticity trained sparse autoencoders on 8 billion activations of the 512-neuron MLP layer of a one-layer Transformer, with dictionaries from 512 to 131,072 features, and studied a 4,096-feature run in detail; features such as Arabic script, DNA sequences, base64 and Hebrew were much more interpretable than the neurons. Established Cunningham and colleagues independently found sparse-autoencoder features more interpretable than other decompositions by automated measures. Established Scaling Monosemanticity trained dictionaries of about 1M, 4M and 34M features on the middle-layer residual stream of Claude 3 Sonnet, finding abstract, multilingual and multimodal features, including safety-relevant ones such as features related to deception and bias. Established

Testing that features are used

A feature that fires on Golden Gate Bridge text could be a correlate the model ignores. Clamping that feature to 10 times its maximum activation during the forward pass made Claude 3 Sonnet start to identify itself as the Golden Gate Bridge. Established Intervening and watching behaviour change is the same logic as causal tests of probes and circuits.

Limits

Dictionaries do not reconstruct activations perfectly, so some of the model's computation lives in the error term; many concepts have no clean feature at a given dictionary size; features split into finer ones as the dictionary grows; and their labels are human interpretations of activation examples. Whether sparse autoencoders find the model's "true" features is an open question. Active research

Tiny example

Take the pentagon from the superposition lab: two hidden numbers, five feature directions. Given a hidden vector that is exactly feature 3's direction, any two directions could reconstruct it, but the sparsest explanation is "feature 3 alone". That preference for the sparsest explanation, at scale, is what the L1 penalty encodes.

Mini experiment

Open the interactive feature browser linked from Towards Monosemanticity and pick three features at random. For each, write your own label from the top activating examples before reading the authors'. How often do you agree, and what would you need to see to trust a label?

What to remember

  • SAE: activation x → wide sparse code f → reconstruction x̂ = b + Σ fᵢ dᵢ. Loss = reconstruction error + λ·Σ|fᵢ|.
  • Each decoder vector dᵢ is a feature direction; read it by the inputs that activate it most.
  • Towards Monosemanticity (2023): a 512-neuron layer, dictionaries of 512 to 131,072 features; features for Arabic script, DNA, base64…
  • Scaling Monosemanticity (2024): ~1M, 4M and 34M features on Claude 3 Sonnet's middle layer.
  • Clamping the Golden Gate Bridge feature made the model talk as if it were the bridge: evidence features are used.
  • Limits: reconstruction is imperfect, many concepts lack a clean feature, and labels are human interpretations.

Key papers

Essential

Towards Monosemanticity: Decomposing Language Models With Dictionary Learning

Trenton Bricken, Adly Templeton et al. · 2023 · Transformer Circuits Thread

Showed that a sparse autoencoder can split a small model's polysemantic neurons into thousands of features that each mean one thing.

How to read it: Start with the summary and one detailed feature (the Arabic-script feature); the interactive feature browser linked from the article is the best way to get a feel for it.

~1 h 30 min read✓ verified 2026-10-07
Important

Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet

Adly Templeton, Tom Conerly et al. · 2024 · Transformer Circuits Thread

Took dictionary learning from a toy-sized model to a production model and showed that clamping a feature changes behaviour.

How to read it: The 'Influence on Behavior' section is the evidence that features are used, not just correlated; the 'Feature Completeness' section is the honest limit (many concepts have no clean feature yet).

~1 h 30 min read✓ verified 2026-10-07