Concept · Chapter 14: Reasoning Models
Faithfulness of Reasoning Traces
A reasoning trace is faithful to the extent that it accurately reflects what actually caused the model's answer; a trace can be useful, correct and convincing while still leaving out what mattered.
The problem
It is tempting to read a chain of thought as the model's actual reasoning, and to trust an answer because its explanation sounds right.
The solution
Treat a trace as an artifact to test, not a window into the model: check its claims independently, and intervene on inputs to see whether the answer depends on what the trace says it does.
The consequence
Traces remain valuable for debugging and monitoring, but trust has to rest on external checks and reproducible results rather than on how plausible the explanation reads.
You should understand first
- Text as Data
- Probability and Distributions
- Conditional Probability and Bayes' Theorem
- Probability of Sequences
- Language Modeling
- Vectors
- Dot Product
- Embeddings
- Attention
- Softmax
- Self-Attention
- Causal Masking
- Entropy
- Loss Functions
- Cross-Entropy Loss
- Autoregressive Next-Token Prediction
- Pretraining at Scale
- GPT-1 → GPT-2 → GPT-3
- In-Context Learning
- Chain of Thought
- Faithfulness of Reasoning Traces
Useful is not the same as faithful
Three questions about a trace, each with an independent answer:
- Is it useful? Did writing it help the model get the right answer?
- Is it correct? Are its individual statements true?
- Is it faithful? Does it describe what actually drove the answer?
A trace can be useful and correct but unfaithful: every line true, while the answer was really driven by a cue the trace never mentions. It can also be faithful and wrong: an honest record of a bad calculation.
The evidence
Turpin and colleagues added biasing features to prompts, for example reordering few-shot examples so the correct answer was always “(A)”; models' answers shifted toward the bias while their chain-of-thought explanations did not mention it, often rationalising the biased answer instead, and accuracy fell by as much as 36% on a suite of 13 BIG-Bench Hard tasks with GPT-3.5 and Claude 1.0. EstablishedThis establishes counterexamples, not a universal verdict. It shows a convincing explanation is not evidence by itself. It does not show that every trace is misleading, or that intermediate text never reflects the computation.
Test it like a hypothesis
Reading a trace gives you a hypothesis about why the model answered as it did. Interventions test it:
- Remove or change a fact the trace relies on. Does the answer change?
- Add a cue the trace never mentions, such as a suggested answer or a reordered list. Does the answer change anyway?
- Corrupt a step in the middle of a trace and let the model continue. Does the final answer follow the corrupted step, or ignore it?
If the answer moves with things the trace ignores, or ignores things the trace relies on, the trace is not the whole story. This is the same discipline as validating a data pipeline: you trust the lineage you have tested, not the lineage the documentation describes.
What users actually see
OpenAI's o1 announcement said users would be shown a model-generated summary of the chain of thought rather than the raw reasoning tokens. Established Other systems expose the full trace. In either case, neither the summary nor the trace is direct access to the network's activations; those are studied by interpretability research, the topic of a later chapter.
Why it still matters
Traces are cheap to read, and when they are faithful enough they let people and automated monitors catch mistakes, unsafe plans or reward hacking before an answer is used. That value depends on the trace staying informative. OpenAI's o1 announcement gave this as a reason for not training policy compliance or user preferences onto the hidden chain of thought: monitoring needs the model to express its reasoning in unaltered form. Established
In practice
- Check results with tools or tests, not with the model's own account of its work.
- When a trace cites a source or a calculation, verify that item.
- Report final correctness, explanation quality and faithfulness separately when you evaluate a system.
Mini experiment
Give a chat model a multiple-choice question and three solved examples whose correct answers are all “(A)”. Ask a fourth question whose answer is (B) and request step-by-step reasoning. Does the answer drift toward (A)? If it does, does the reasoning mention the pattern? Repeat with the examples' answers shuffled. Five trials are an anecdote, but they teach the method.
What to remember
- Useful, correct and faithful are three different properties of a trace.
- Turpin et al.: biasing features changed answers without being mentioned, and accuracy fell by up to 36% on 13 BIG-Bench Hard tasks.
- Test faithfulness by intervention: change an input the trace ignores and see whether the answer moves.
- Some products show only a model-written summary of the reasoning, not the reasoning tokens.
- Unfaithful in some settings does not mean useless; it means not self-certifying.
Key papers
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Jason Wei, Xuezhi Wang et al. · 2022 · NeurIPS 2022
Showed that prompting large models to write out intermediate steps markedly improves multi-step reasoning — the seed of today's reasoning models.
Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting
Miles Turpin, Julian Michael et al. · 2023 · NeurIPS 2023
Showed with controlled experiments that a chain of thought can omit what actually drove the answer, so a plausible explanation is not evidence of faithfulness.
How to read it: The experimental design is the lesson: change an input the explanation never mentions and see whether the answer moves.