Skip to content
Road to Intelligence

Concept · Chapter 16: Evaluation, Reliability, Safety & Interpretability

Circuits and Mechanistic Interpretability

Should knowUnderstand14 minDifficulty

Mechanistic interpretability tries to reverse-engineer the algorithm a network has learned, as circuits of components (attention heads, MLPs, features) that pass information to each other, and tests each claim by intervening on the components.

The problem

Probes and features say what information is present, not how the model turns its input into its output; for that you need the steps of the computation.

The solution

Treat the residual stream as a shared channel that components read from and write to, find candidate components for a behaviour, and confirm them with interventions such as activation patching: replace a component's activation with one from a different input and measure the change in output.

The consequence

Small behaviours in real models have been explained step by step (induction heads, indirect-object identification), and the methods scale slowly; full explanations of large models remain far off.

The residual stream as a bus

Elhage and colleagues described a Transformer's residual stream as a communication channel: each attention head and MLP reads from it and adds its output back, so the components act as independent, additive writers. Established For a data engineer: a shared message bus where every stage reads the current record, appends a field and passes it on. A circuit is a chain of such stages that implements one behaviour.

Induction heads

The cleanest example. In small attention-only Transformers, the same authors found that two-layer models compose heads into "induction heads", which only appear with at least two attention layers and explain in-context learning in those models. Established

It takes two heads:

  1. A previous-token head in the first layer writes, at each position, "the token before me was …".
  2. An induction head in the second layer, at the current token [A], looks for earlier positions whose previous token was [A], attends to them, and copies what is there: [B].

So "Mr Dursley … Mr" predicts "Dursley". Olsson and colleagues linked induction heads to in-context learning in larger models with several lines of evidence, causal in small models and correlational in larger ones. Established

Activation patching

Finding components is half the job; showing they matter is the other. Meng and colleagues ran the model on a prompt such as "The Space Needle is in downtown" and on a corrupted copy with the subject obscured, then restored individual internal activations from the clean run; restoring mid-layer MLP activations at the subject's last token brought the right answer back. Established They then edited single facts by a rank-one update to one MLP's weights. Established Hase and colleagues later found that where causal tracing localized a fact did not predict which layer was best to edit. Established Localizing and controlling are different questions.

A circuit in the wild

Wang and colleagues explained how GPT-2 small completes "When Mary and John went to the store, John gave a drink to" with "Mary": 26 attention heads in 7 classes, found with causal interventions and evaluated for faithfulness, completeness and minimality, criteria that also exposed remaining gaps. Established In a small Transformer trained on modular addition, Nanda and colleagues reverse-engineered the complete algorithm, which uses discrete Fourier transforms and trigonometric identities. Established

Tiny example

Sequence: "A7 B3 … A7". At the last token, the induction head's query is "find positions whose previous token was A7". Only the position of "B3" qualifies (its previous token is A7), so the head attends there and copies "B3" into the stream, raising its logit. Patch in the first layer's output from a sequence where the earlier pair was "A7 C9", and the prediction changes to "C9": evidence the head is copying, not guessing.

Mini experiment

Write a 12-token sequence with one repeated pair, such as "red 4 blue 9 green 2 red". By hand, mark which earlier position an induction head should attend to at the last token and what it predicts. Then write one where the pattern is ambiguous (two earlier "red"s followed by different tokens). What should the head do, and what would you measure to find out?

What to remember

  • Residual stream: every layer reads from and adds to it; heads and MLPs are independent writers.
  • Induction head: at [A][B] … [A], attend to the token after the earlier [A] and copy [B]. Needs two attention layers (a previous-token head + the induction head).
  • Activation patching / causal tracing: run a clean and a corrupted input, restore one activation, see if the answer comes back.
  • ROME (2022): factual recall traced to mid-layer MLPs at the subject's last token; but later work found tracing didn't predict the best layer to edit.
  • IOI circuit (2022): 26 heads in 7 classes in GPT-2 small, scored for faithfulness, completeness, minimality.
  • Explanations are claims to be tested, and current ones cover small behaviours.

Key papers

Optional

In-context Learning and Induction Heads

Catherine Olsson, Nelson Elhage et al. · 2022

Proposed a concrete mechanism: 'induction heads', attention heads that complete [A][B] … [A] → [B]. Their appearance during training coincides with a jump in in-context learning.

~1 h readarXiv:2209.11895✓ verified 2026-09-26
Important

Locating and Editing Factual Associations in GPT

Kevin Meng, David Bau et al. · 2022

Made activation patching ('causal tracing') a standard tool, and tested a localization claim by editing the weights it pointed to.

How to read it: Figure 1 explains causal tracing in one diagram. Then read Hase et al. (2023), which found that tracing results did not predict which layer is best to edit: localization and editing are separate questions.

~40 min readarXiv:2202.05262✓ verified 2026-10-07
Important

A Mathematical Framework for Transformer Circuits

Nelson Elhage, Neel Nanda et al. · 2021 · Transformer Circuits Thread

The vocabulary of mechanistic interpretability for Transformers: the residual stream as a shared channel, heads as independent readers and writers, and induction heads.

How to read it: Read the summary of results and the 'Induction Heads' section; the path-expansion algebra in between is worth it only if you plan to do this work.

~1 h 30 min read✓ verified 2026-10-07