Skip to content
Road to Intelligence

Concept · Chapter 14: Reasoning Models

Faithfulness of Reasoning Traces

Should knowKnow well10 minDifficulty

A reasoning trace is faithful to the extent that it accurately reflects what actually caused the model's answer; a trace can be useful, correct and convincing while still leaving out what mattered.

The problem

It is tempting to read a chain of thought as the model's actual reasoning, and to trust an answer because its explanation sounds right.

The solution

Treat a trace as an artifact to test, not a window into the model: check its claims independently, and intervene on inputs to see whether the answer depends on what the trace says it does.

The consequence

Traces remain valuable for debugging and monitoring, but trust has to rest on external checks and reproducible results rather than on how plausible the explanation reads.

Useful is not the same as faithful

Three questions about a trace, each with an independent answer:

  • Is it useful? Did writing it help the model get the right answer?
  • Is it correct? Are its individual statements true?
  • Is it faithful? Does it describe what actually drove the answer?

A trace can be useful and correct but unfaithful: every line true, while the answer was really driven by a cue the trace never mentions. It can also be faithful and wrong: an honest record of a bad calculation.

The evidence

Turpin and colleagues added biasing features to prompts, for example reordering few-shot examples so the correct answer was always “(A)”; models' answers shifted toward the bias while their chain-of-thought explanations did not mention it, often rationalising the biased answer instead, and accuracy fell by as much as 36% on a suite of 13 BIG-Bench Hard tasks with GPT-3.5 and Claude 1.0. Established

This establishes counterexamples, not a universal verdict. It shows a convincing explanation is not evidence by itself. It does not show that every trace is misleading, or that intermediate text never reflects the computation.

Test it like a hypothesis

Reading a trace gives you a hypothesis about why the model answered as it did. Interventions test it:

  • Remove or change a fact the trace relies on. Does the answer change?
  • Add a cue the trace never mentions, such as a suggested answer or a reordered list. Does the answer change anyway?
  • Corrupt a step in the middle of a trace and let the model continue. Does the final answer follow the corrupted step, or ignore it?

If the answer moves with things the trace ignores, or ignores things the trace relies on, the trace is not the whole story. This is the same discipline as validating a data pipeline: you trust the lineage you have tested, not the lineage the documentation describes.

What users actually see

OpenAI's o1 announcement said users would be shown a model-generated summary of the chain of thought rather than the raw reasoning tokens. Established Other systems expose the full trace. In either case, neither the summary nor the trace is direct access to the network's activations; those are studied by interpretability research, the topic of a later chapter.

Why it still matters

Traces are cheap to read, and when they are faithful enough they let people and automated monitors catch mistakes, unsafe plans or reward hacking before an answer is used. That value depends on the trace staying informative. OpenAI's o1 announcement gave this as a reason for not training policy compliance or user preferences onto the hidden chain of thought: monitoring needs the model to express its reasoning in unaltered form. Established

In practice

  • Check results with tools or tests, not with the model's own account of its work.
  • When a trace cites a source or a calculation, verify that item.
  • Report final correctness, explanation quality and faithfulness separately when you evaluate a system.

Mini experiment

Give a chat model a multiple-choice question and three solved examples whose correct answers are all “(A)”. Ask a fourth question whose answer is (B) and request step-by-step reasoning. Does the answer drift toward (A)? If it does, does the reasoning mention the pattern? Repeat with the examples' answers shuffled. Five trials are an anecdote, but they teach the method.

What to remember

  • Useful, correct and faithful are three different properties of a trace.
  • Turpin et al.: biasing features changed answers without being mentioned, and accuracy fell by up to 36% on 13 BIG-Bench Hard tasks.
  • Test faithfulness by intervention: change an input the trace ignores and see whether the answer moves.
  • Some products show only a model-written summary of the reasoning, not the reasoning tokens.
  • Unfaithful in some settings does not mean useless; it means not self-certifying.

Key papers

Important

Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting

Miles Turpin, Julian Michael et al. · 2023 · NeurIPS 2023

Showed with controlled experiments that a chain of thought can omit what actually drove the answer, so a plausible explanation is not evidence of faithfulness.

How to read it: The experimental design is the lesson: change an input the explanation never mentions and see whether the answer moves.

~35 min readarXiv:2305.04388✓ verified 2026-10-06