Skip to content
Road to Intelligence

Concept · Chapter 16: Evaluation, Reliability, Safety & Interpretability

Probing Representations

Should knowKnow well11 minDifficulty

A probe is a small classifier, usually linear, trained to predict some property from a model's internal activations, which shows what information is available at that layer but not, by itself, that the model uses it.

The problem

A network's activations are thousands of numbers per token with no labels, so it is unclear what each layer represents or where a property such as part of speech or truth becomes available.

The solution

Collect activations for inputs whose property you know, train a simple classifier on them, and compare its accuracy with careful controls; then test causally by changing the activations along the probe's direction.

The consequence

Probes are the quickest way to ask 'is X in here?', and they come with two standard cautions: a powerful probe can learn the task itself, and decodable is not the same as used.

Ask the activations a question

Run a model on many inputs whose property you know (say, whether a statement is true) and save the activation vector at some layer for each. Train a logistic regression from those vectors to the property. If it predicts well on held-out inputs, the information is linearly available at that layer.

Alain and Bengio proposed such linear classifier "probes", trained independently of the model, and found in Inception v3 and ResNet-50 that linear separability of the classes increased monotonically with depth. Established That is the representation learning story of Chapter 4, measured.

Two cautions

The probe might do the work. Hewitt and Liang introduced control tasks, which assign random labels to word types so that only the probe itself can learn them, and found that popular probes on ELMo scored well on these too; they call a probe selective when it does well on the real task and poorly on the control. Established A large enough probe can memorize its way to high accuracy from almost any representation. Keep probes simple and report a control.

Available is not used. A property can be decodable at a layer without the model relying on it. The stronger test is causal: Marks and Tegmark found linear structure separating true from false statements, probes that transferred across datasets, and interventions along the probe direction that made the model treat false statements as true and vice versa. Established

Without labels

Burns and colleagues found a direction for yes/no questions with no labels at all, by requiring that a statement and its negation get opposite truth values; it beat zero-shot accuracy by 4% on average across 6 models and 10 datasets and stayed accurate when the model was prompted to answer wrongly. Established Whether such directions reflect what a model "believes", and how reliably they generalize, is still debated. Active research

Tiny example

Suppose each activation is two numbers and true statements cluster around (2, 1), false ones around (−1, 0). The direction from the false mean to the true mean is (3, 1); projecting any activation on it and thresholding is a probe ("difference in means"). To test it causally, add a multiple of (3, 1) to a false statement's activation inside the model and see whether its next words treat the statement as true.

Mini experiment

Using any open model and a library such as scikit-learn, collect last-token activations for 200 simple true and false statements ("Paris is in France", "Paris is in Peru"). Train a logistic regression on 150, test on 50. Then shuffle the labels and train again. The gap between the two accuracies is your evidence.

What to remember

  • Probe = classifier from activations at a layer to a property; keep it simple (linear) so it reads rather than computes.
  • Alain & Bengio 2016: linear separability of classes increased with depth in Inception v3 and ResNet-50.
  • Control tasks (Hewitt & Liang 2019): compare with random labels; a good probe is selective.
  • Decodable ≠ used: test by intervening (move activations along the direction and watch the output).
  • Truth probes: simple true/false statements appear linearly separable in large models, with causal evidence (Marks & Tegmark 2023).

Key papers

Optional

Designing and Interpreting Probes with Control Tasks

John Hewitt, Percy Liang · 2019

Asked the question every probing result must answer: did the representation encode the property, or did the probe learn it?

~25 min readarXiv:1909.03368✓ verified 2026-10-07