Skip to content
Road to Intelligence

Concept · Chapter 16: Evaluation, Reliability, Safety & Interpretability

Superposition and Polysemantic Neurons

Should knowKnow well15 minDifficulty

Superposition is a network storing more features than it has neurons, as overlapping directions in activation space, which works when features are rarely active at the same time and explains why single neurons often respond to unrelated things.

The problem

Interpretability would be easy if each neuron meant one thing, but many neurons in real models fire for several unrelated concepts, so reading a network neuron by neuron fails.

The solution

Model features as directions rather than neurons: when features are sparse, a network can pack many of them into fewer dimensions at small angles and filter the resulting interference with nonlinearities and biases. Recover the features with methods that look for directions, such as sparse autoencoders.

The consequence

The natural units of a network are directions, not neurons; this reframed mechanistic interpretability around finding features first and led directly to dictionary learning on large models.

Neurons that mean several things

If every neuron in a model detected one concept, you could read the model by listing what each neuron responds to. Real models disappoint. In the small Transformer studied in Towards Monosemanticity, a single neuron responds to a mixture of academic citations, English dialogue, HTTP requests and Korean text. Established Neurons like this are polysemantic. As early as 2013, Szegedy and colleagues found no distinction between individual high-level units and random linear combinations of them under the unit-analysis methods they tried, suggesting that the space, not the units, carries the meaning. Established

More features than dimensions

A model may need to track far more features than it has neurons: every name, topic, language, syntactic role. In mm dimensions you can fit only mm exactly perpendicular directions, but you can fit many more that are almost perpendicular. If two features are rarely active together, giving them overlapping directions costs little.

Elhage and colleagues demonstrated this with toy models: h=Wxh = Wx compresses nn sparse features into m<nm < n dimensions and x′=ReLU(W⊤Wx+b)x' = \mathrm{ReLU}(W^\top W x + b) tries to reconstruct them. With dense features, the model learns an orthogonal basis for the most important features, like PCA, and drops the rest; as features become sparse, it represents more of them in superposition, first as antipodal pairs and then in geometric arrangements such as pentagons. Established

Why the ReLU matters

With overlapping directions, switching on feature 1 also nudges the outputs of its neighbours: interference. The paper finds that a linear model never uses superposition, while the ReLU output model does, using a negative bias so that small interference falls below zero and is removed. Established The model accepts a little error when two features happen to fire together, in exchange for representing many more features the rest of the time.

Tiny example (lab numbers)

In the lab, five features with importances 1, 0.7, 0.49, 0.34 and 0.24 must pass through two hidden numbers.

  • Dense (every feature active on every input): after training, features 1 and 2 have length 1.0 and sit 90° apart; features 3–5 have length about 0 and their biases settle near 0.5, the average value of a feature. Two features, two dimensions.
  • 30% density: four features, as two antipodal pairs (opposite directions), at 90° to each other. The least important feature is dropped.
  • 3% density: all five, at roughly 72° apart, a pentagon, with biases near −0.25 to filter interference. (With decaying importances this happens on 6 of 10 random seeds; the other 4 keep four features. With equal importances, 10 of 10 form a pentagon.)

In real models

Anthropic's Towards Monosemanticity decomposed the 512-neuron MLP layer of a one-layer Transformer into thousands of features with sparse autoencoders, finding features far more interpretable than the neurons, consistent with superposition. Established The picture that results: a layer's activation is a sum of a few active feature directions out of a very large dictionary, and a neuron is just one coordinate of that sum. Interpretation Sparse autoencoders are the tool for recovering the dictionary.

Mini experiment

In the lab, train at 3% density and click a feature to switch it on alone. Which other outputs move, and why don't the ones on the far side of the pentagon? Then train at 100% density and repeat. Where did features 3–5 go?

Why should I care?

As a researcher

Superposition is why interpretability needs dictionary learning and why 'what does this neuron do?' is often the wrong question; it is also linked, in the original paper, to adversarial examples.

As an engineer

It explains why steering or editing a model by touching single neurons has side effects: the same neuron carries other features.

Modern systems that depend on it

  • sparse autoencoders
  • feature steering
  • circuit analysis with features
  • the linear representation hypothesis

Historical context

Before

Interpretability looked for neurons with clean meanings, and found many but also many that responded to several unrelated inputs.

After

Features are treated as directions possibly shared across many neurons; toy models predict when superposition happens, and dictionary learning extracts features from real models.

Used today

The basis for sparse-autoencoder feature dictionaries, feature-level circuit analysis, and feature steering experiments on production models.

What to remember

  • Polysemantic neuron: responds to several unrelated things. Superposition is one explanation.
  • Toy model: h = Wx with 2 hidden numbers for 5 features, output ReLU(WᵀWx + b).
  • Dense features → keep the 2 most important, at right angles (like PCA); the rest set to their average.
  • Sparse features → store more than 2: antipodal pairs, then a pentagon (5 directions 72° apart).
  • Interference is filtered by the ReLU and a negative bias; it is cheap because features rarely co-occur.
  • Feature = direction in activation space, not necessarily a neuron.

Key papers

Important

Intriguing properties of neural networks

Christian Szegedy, Wojciech Zaremba et al. · 2013

Introduced adversarial examples, and in the same paper observed that meaning lives in directions of activation space rather than single units: two threads that run through this chapter.

~25 min readarXiv:1312.6199✓ verified 2026-10-07
Essential

Toy Models of Superposition

Nelson Elhage, Tristan Hume et al. · 2022

Gave the leading explanation for why single neurons are hard to interpret: networks store more features than they have dimensions.

How to read it: Read up to 'Mathematical Understanding' with the figures; the lab in this chapter reproduces the five-features-in-two-dimensions example. The later sections on phase diagrams and geometry are optional.

~1 h readarXiv:2209.10652✓ verified 2026-10-07
Essential

Towards Monosemanticity: Decomposing Language Models With Dictionary Learning

Trenton Bricken, Adly Templeton et al. · 2023 · Transformer Circuits Thread

Showed that a sparse autoencoder can split a small model's polysemantic neurons into thousands of features that each mean one thing.

How to read it: Start with the summary and one detailed feature (the Arabic-script feature); the interactive feature browser linked from the article is the best way to get a feel for it.

~1 h 30 min read✓ verified 2026-10-07

Watch