Concept · Chapter 14: Reasoning Models
Chain of Thought
A chain of thought is intermediate text generated before the final answer, so that later tokens can condition on partial results instead of computing everything implicitly in one step.
The problem
Asked for the answer directly, a model must produce the result's first token with all the intermediate arithmetic or logic still implicit in its activations, and multi-step problems often fail.
The solution
Let the model write intermediate steps first, by showing worked examples, by instructing it to reason step by step, or by training it to do so; each written result becomes context for the next step.
The consequence
Multi-step accuracy rose sharply on many benchmarks, and chains of thought became the substrate that later methods sample, search, score and reward, while raising a new question: does the written chain reflect how the answer was produced?
You should understand first
- Text as Data
- Probability and Distributions
- Conditional Probability and Bayes' Theorem
- Probability of Sequences
- Language Modeling
- Vectors
- Dot Product
- Embeddings
- Attention
- Softmax
- Self-Attention
- Causal Masking
- Entropy
- Loss Functions
- Cross-Entropy Loss
- Autoregressive Next-Token Prediction
- Pretraining at Scale
- GPT-1 → GPT-2 → GPT-3
- In-Context Learning
- Chain of Thought
Put a result where the next step can use it
Six trays hold seven seedlings each, and half are moved outside. How many remain?
Answering “21” immediately means the model must get from the question to the first token of the answer with and never written anywhere. Writing “6 × 7 = 42” first changes the situation: the number 42 is now a token in the context, and the next prediction can attend to it like any other input. Each written step is a small, easier prediction conditioned on the previous ones. (This example is written for teaching; it is not a captured model trace.)
A data-pipeline analogy: computing everything in one giant query versus materialising intermediate tables. The intermediate tables cost storage, but each later step becomes simpler and inspectable.
Three ways to get one
Wei and colleagues showed that few-shot prompts containing worked, step-by-step examples improved large models' performance on arithmetic, commonsense and symbolic reasoning tasks in their experiments; the gains appeared mainly in sufficiently large models. Established Kojima and colleagues then found that no examples were needed: adding “Let's think step by step” raised InstructGPT's accuracy on MultiArith from 17.7% to 78.7% and on GSM8K from 10.4% to 40.7%. EstablishedKeep three mechanisms apart, because they produce similar-looking text:
- Prompting (worked examples or an instruction) changes the input. Weights are untouched.
- Supervised fine-tuning on worked solutions changes weights to imitate the solutions.
- Reinforcement learning changes weights according to a reward on the outcome or the steps, see verifiable rewards.
Today's reasoning models mostly come from the third route, starting from a model that already writes chains of thought.
The probability picture
Write for the question, for the intermediate text and for the answer. The model's overall probability of an answer sums over all the chains it might write:
What it does. One generated response samples (or greedily decodes) a single , then an answer conditioned on it. It does not compute the sum. If the model writes a bad chain, the answer conditions on the bad chain. That is why sampling several chains and voting, self-consistency, approximates the sum better than one chain does.
Costs and failure modes
- Tokens and latency. Every intermediate token is generated sequentially and billed. A chain of 2,000 tokens takes longer than a 20-token answer, whatever the hardware.
- Error propagation. A wrong step early in the chain becomes context for everything after it.
- Padding. Text such as “let me think carefully” adds tokens without adding a partial result. The reward lab shows what happens if length itself is rewarded.
- Parsing. The final answer must be extracted from free text, which is a source of evaluation errors.
- Faithfulness. Turpin and colleagues showed that biasing features in a prompt can change a model's answer without appearing in its chain of thought. Established See reasoning faithfulness.
Mini experiment
Ask any chat model a two-step word problem twice: once with “Answer with only the number” and once with “Show each calculation, then give the number”. Repeat on five problems. Count correct answers and output tokens for each mode. You have measured, on a tiny sample, the trade this page describes: tokens for accuracy. Treat five problems as an anecdote, not a benchmark.
Why should I care?
As a researcher
Nearly every reasoning method in the literature operates on chains of thought: they are sampled, voted on, searched over, scored step by step and rewarded.
As an engineer
Asking for intermediate work is the cheapest reasoning upgrade available, and it changes latency, token cost and how you must parse the final answer.
Modern systems that depend on it
- self-consistency
- tree search over thoughts
- process reward models
- reasoning models trained with RL
Historical context
Before
Prompts asked for the answer directly, and models that failed multi-step problems seemed to lack the capability altogether.
After
Models write intermediate steps before answering, by prompting at first and later by training, and the steps themselves became an object to sample, check and reward.
Used today
Reasoning models are trained to produce long intermediate work before answering; some products show it, others show only a summary.
What to remember
- Intermediate text puts partial results into the context, where later tokens can use them.
- Wei et al. (2022) used worked examples; Kojima et al. (2022) got large gains from “Let's think step by step” alone.
- Prompting, fine-tuning on worked solutions and RL are three different ways to get a chain of thought; only the last two change weights.
- Longer is not better: a chain can carry an early mistake forward, or pad without computing anything.
- A plausible chain is not proof of how the answer was produced; check the result independently.
Key papers
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Jason Wei, Xuezhi Wang et al. · 2022 · NeurIPS 2022
Showed that prompting large models to write out intermediate steps markedly improves multi-step reasoning — the seed of today's reasoning models.
Large Language Models are Zero-Shot Reasoners
Takeshi Kojima, Shixiang Shane Gu et al. · 2022 · NeurIPS 2022
Showed that chain-of-thought behaviour needs no worked examples: one generic instruction elicits it, which made step-by-step prompting a default habit.
How to read it: The two-stage prompt (first elicit reasoning, then extract the answer) is a detail worth noticing: the answer still has to be parsed from free text.
Watch
Andrej Karpathy
Deep Dive into LLMs like ChatGPT
A long, general-audience walk through the whole pipeline behind a chat model, from internet text to tokens to pretraining to post-training.