Concept · Chapter 16: Evaluation, Reliability, Safety & Interpretability
Probing Representations
A probe is a small classifier, usually linear, trained to predict some property from a model's internal activations, which shows what information is available at that layer but not, by itself, that the model uses it.
The problem
A network's activations are thousands of numbers per token with no labels, so it is unclear what each layer represents or where a property such as part of speech or truth becomes available.
The solution
Collect activations for inputs whose property you know, train a simple classifier on them, and compare its accuracy with careful controls; then test causally by changing the activations along the probe's direction.
The consequence
Probes are the quickest way to ask 'is X in here?', and they come with two standard cautions: a powerful probe can learn the task itself, and decodable is not the same as used.
You should understand first
- The Turing Test
- Symbolic AI
- Logic and Rules
- Expert Systems
- Knowledge Representation
- The Knowledge-Acquisition Bottleneck
- From Rules to Learning
- Supervised, Unsupervised and Self-Supervised Learning
- Vectors
- Features, Labels and Tasks
- Loss Functions
- Derivatives and Gradients
- Gradient Descent
- Linear Regression
- Probability and Distributions
- Entropy
- Softmax
- Cross-Entropy Loss
- Logistic Regression
- Hand-Crafted Features vs Learned Features
- Dot Product
- The Perceptron
- Activation Functions
- The Artificial Neuron
- Matrix Multiplication
- Multilayer Perceptron (MLP)
- Representation Learning
- Embeddings
- Probing Representations
Ask the activations a question
Run a model on many inputs whose property you know (say, whether a statement is true) and save the activation vector at some layer for each. Train a logistic regression from those vectors to the property. If it predicts well on held-out inputs, the information is linearly available at that layer.
Alain and Bengio proposed such linear classifier "probes", trained independently of the model, and found in Inception v3 and ResNet-50 that linear separability of the classes increased monotonically with depth. Established That is the representation learning story of Chapter 4, measured.
Two cautions
The probe might do the work. Hewitt and Liang introduced control tasks, which assign random labels to word types so that only the probe itself can learn them, and found that popular probes on ELMo scored well on these too; they call a probe selective when it does well on the real task and poorly on the control. Established A large enough probe can memorize its way to high accuracy from almost any representation. Keep probes simple and report a control.
Available is not used. A property can be decodable at a layer without the model relying on it. The stronger test is causal: Marks and Tegmark found linear structure separating true from false statements, probes that transferred across datasets, and interventions along the probe direction that made the model treat false statements as true and vice versa. Established
Without labels
Burns and colleagues found a direction for yes/no questions with no labels at all, by requiring that a statement and its negation get opposite truth values; it beat zero-shot accuracy by 4% on average across 6 models and 10 datasets and stayed accurate when the model was prompted to answer wrongly. Established Whether such directions reflect what a model "believes", and how reliably they generalize, is still debated. Active researchTiny example
Suppose each activation is two numbers and true statements cluster around (2, 1), false ones around (−1, 0). The direction from the false mean to the true mean is (3, 1); projecting any activation on it and thresholding is a probe ("difference in means"). To test it causally, add a multiple of (3, 1) to a false statement's activation inside the model and see whether its next words treat the statement as true.
Mini experiment
Using any open model and a library such as scikit-learn, collect last-token activations for 200 simple true and false statements ("Paris is in France", "Paris is in Peru"). Train a logistic regression on 150, test on 50. Then shuffle the labels and train again. The gap between the two accuracies is your evidence.
What to remember
- Probe = classifier from activations at a layer to a property; keep it simple (linear) so it reads rather than computes.
- Alain & Bengio 2016: linear separability of classes increased with depth in Inception v3 and ResNet-50.
- Control tasks (Hewitt & Liang 2019): compare with random labels; a good probe is selective.
- Decodable ≠ used: test by intervening (move activations along the direction and watch the output).
- Truth probes: simple true/false statements appear linearly separable in large models, with causal evidence (Marks & Tegmark 2023).
Key papers
Understanding intermediate layers using linear classifier probes
Guillaume Alain, Yoshua Bengio · 2016
Named and popularised the linear probe, the simplest tool for asking what information a layer contains.
Designing and Interpreting Probes with Control Tasks
John Hewitt, Percy Liang · 2019
Asked the question every probing result must answer: did the representation encode the property, or did the probe learn it?
The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets
Samuel Marks, Max Tegmark · 2023
Careful evidence that models represent whether a simple statement is true along a direction you can find, transfer and intervene on.
Discovering Latent Knowledge in Language Models Without Supervision
Collin Burns, Haotian Ye et al. · 2022
Asked whether we can read what a model 'believes' from its activations, separately from what it says.