Concept · Chapter 16: Evaluation, Reliability, Safety & Interpretability
Circuits and Mechanistic Interpretability
Mechanistic interpretability tries to reverse-engineer the algorithm a network has learned, as circuits of components (attention heads, MLPs, features) that pass information to each other, and tests each claim by intervening on the components.
The problem
Probes and features say what information is present, not how the model turns its input into its output; for that you need the steps of the computation.
The solution
Treat the residual stream as a shared channel that components read from and write to, find candidate components for a behaviour, and confirm them with interventions such as activation patching: replace a component's activation with one from a different input and measure the change in output.
The consequence
Small behaviours in real models have been explained step by step (induction heads, indirect-object identification), and the methods scale slowly; full explanations of large models remain far off.
You should understand first
- Vectors
- Dot Product
- Embeddings
- Attention
- Probability and Distributions
- Softmax
- Self-Attention
- Multi-Head Attention
- Residual Connections
- Text as Data
- Conditional Probability and Bayes' Theorem
- Probability of Sequences
- Language Modeling
- Causal Masking
- Entropy
- Loss Functions
- Cross-Entropy Loss
- Autoregressive Next-Token Prediction
- Pretraining at Scale
- GPT-1 → GPT-2 → GPT-3
- In-Context Learning
- Circuits and Mechanistic Interpretability
The residual stream as a bus
Elhage and colleagues described a Transformer's residual stream as a communication channel: each attention head and MLP reads from it and adds its output back, so the components act as independent, additive writers. Established For a data engineer: a shared message bus where every stage reads the current record, appends a field and passes it on. A circuit is a chain of such stages that implements one behaviour.
Induction heads
The cleanest example. In small attention-only Transformers, the same authors found that two-layer models compose heads into "induction heads", which only appear with at least two attention layers and explain in-context learning in those models. Established
It takes two heads:
- A previous-token head in the first layer writes, at each position, "the token before me was …".
- An induction head in the second layer, at the current token [A], looks for earlier positions whose previous token was [A], attends to them, and copies what is there: [B].
So "Mr Dursley … Mr" predicts "Dursley". Olsson and colleagues linked induction heads to in-context learning in larger models with several lines of evidence, causal in small models and correlational in larger ones. Established
Activation patching
Finding components is half the job; showing they matter is the other. Meng and colleagues ran the model on a prompt such as "The Space Needle is in downtown" and on a corrupted copy with the subject obscured, then restored individual internal activations from the clean run; restoring mid-layer MLP activations at the subject's last token brought the right answer back. Established They then edited single facts by a rank-one update to one MLP's weights. Established Hase and colleagues later found that where causal tracing localized a fact did not predict which layer was best to edit. Established Localizing and controlling are different questions.
A circuit in the wild
Wang and colleagues explained how GPT-2 small completes "When Mary and John went to the store, John gave a drink to" with "Mary": 26 attention heads in 7 classes, found with causal interventions and evaluated for faithfulness, completeness and minimality, criteria that also exposed remaining gaps. Established In a small Transformer trained on modular addition, Nanda and colleagues reverse-engineered the complete algorithm, which uses discrete Fourier transforms and trigonometric identities. EstablishedTiny example
Sequence: "A7 B3 … A7". At the last token, the induction head's query is "find positions whose previous token was A7". Only the position of "B3" qualifies (its previous token is A7), so the head attends there and copies "B3" into the stream, raising its logit. Patch in the first layer's output from a sequence where the earlier pair was "A7 C9", and the prediction changes to "C9": evidence the head is copying, not guessing.
Mini experiment
Write a 12-token sequence with one repeated pair, such as "red 4 blue 9 green 2 red". By hand, mark which earlier position an induction head should attend to at the last token and what it predicts. Then write one where the pattern is ambiguous (two earlier "red"s followed by different tokens). What should the head do, and what would you measure to find out?
What to remember
- Residual stream: every layer reads from and adds to it; heads and MLPs are independent writers.
- Induction head: at [A][B] … [A], attend to the token after the earlier [A] and copy [B]. Needs two attention layers (a previous-token head + the induction head).
- Activation patching / causal tracing: run a clean and a corrupted input, restore one activation, see if the answer comes back.
- ROME (2022): factual recall traced to mid-layer MLPs at the subject's last token; but later work found tracing didn't predict the best layer to edit.
- IOI circuit (2022): 26 heads in 7 classes in GPT-2 small, scored for faithfulness, completeness, minimality.
- Explanations are claims to be tested, and current ones cover small behaviours.
Key papers
Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?
Sewon Min, Xinxi Lyu et al. · 2022 · EMNLP 2022
A surprising result: replacing the labels in few-shot examples with random ones barely hurt performance. The examples mainly showed the format, the label space and the kind of input.
In-context Learning and Induction Heads
Catherine Olsson, Nelson Elhage et al. · 2022
Proposed a concrete mechanism: 'induction heads', attention heads that complete [A][B] … [A] → [B]. Their appearance during training coincides with a jump in in-context learning.
Locating and Editing Factual Associations in GPT
Kevin Meng, David Bau et al. · 2022
Made activation patching ('causal tracing') a standard tool, and tested a localization claim by editing the weights it pointed to.
How to read it: Figure 1 explains causal tracing in one diagram. Then read Hase et al. (2023), which found that tracing results did not predict which layer is best to edit: localization and editing are separate questions.
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Kevin Wang, Alexandre Variengien et al. · 2022
The first large end-to-end circuit found in a real language model, with explicit tests of how good the explanation is.
Does Localization Inform Editing? Surprising Differences in Causality-Based Localization vs. Knowledge Editing in Language Models
Peter Hase, Mohit Bansal et al. · 2023
A useful check on interpretability claims: finding where a computation happens does not automatically tell you where to intervene.
Progress measures for grokking via mechanistic interpretability
Neel Nanda, Lawrence Chan et al. · 2023
A complete reverse-engineering of a small trained network, used to explain a training phenomenon that looked sudden.
A Mathematical Framework for Transformer Circuits
Nelson Elhage, Neel Nanda et al. · 2021 · Transformer Circuits Thread
The vocabulary of mechanistic interpretability for Transformers: the residual stream as a shared channel, heads as independent readers and writers, and induction heads.
How to read it: Read the summary of results and the 'Induction Heads' section; the path-expansion algebra in between is worth it only if you plan to do this work.