Skip to content
Road to Intelligence

Concept · Chapter 14: Reasoning Models

Chain of Thought

Must knowKnow well11 minDifficulty

A chain of thought is intermediate text generated before the final answer, so that later tokens can condition on partial results instead of computing everything implicitly in one step.

The problem

Asked for the answer directly, a model must produce the result's first token with all the intermediate arithmetic or logic still implicit in its activations, and multi-step problems often fail.

The solution

Let the model write intermediate steps first, by showing worked examples, by instructing it to reason step by step, or by training it to do so; each written result becomes context for the next step.

The consequence

Multi-step accuracy rose sharply on many benchmarks, and chains of thought became the substrate that later methods sample, search, score and reward, while raising a new question: does the written chain reflect how the answer was produced?

Put a result where the next step can use it

Six trays hold seven seedlings each, and half are moved outside. How many remain?

Answering “21” immediately means the model must get from the question to the first token of the answer with 6×7=426 \times 7 = 42 and 42÷242 \div 2 never written anywhere. Writing “6 × 7 = 42” first changes the situation: the number 42 is now a token in the context, and the next prediction can attend to it like any other input. Each written step is a small, easier prediction conditioned on the previous ones. (This example is written for teaching; it is not a captured model trace.)

A data-pipeline analogy: computing everything in one giant query versus materialising intermediate tables. The intermediate tables cost storage, but each later step becomes simpler and inspectable.

Three ways to get one

Wei and colleagues showed that few-shot prompts containing worked, step-by-step examples improved large models' performance on arithmetic, commonsense and symbolic reasoning tasks in their experiments; the gains appeared mainly in sufficiently large models. Established Kojima and colleagues then found that no examples were needed: adding “Let's think step by step” raised InstructGPT's accuracy on MultiArith from 17.7% to 78.7% and on GSM8K from 10.4% to 40.7%. Established

Keep three mechanisms apart, because they produce similar-looking text:

  1. Prompting (worked examples or an instruction) changes the input. Weights are untouched.
  2. Supervised fine-tuning on worked solutions changes weights to imitate the solutions.
  3. Reinforcement learning changes weights according to a reward on the outcome or the steps, see verifiable rewards.

Today's reasoning models mostly come from the third route, starting from a model that already writes chains of thought.

The probability picture

Write xx for the question, zz for the intermediate text and yy for the answer. The model's overall probability of an answer sums over all the chains it might write:

p(y∣x)=∑zp(z∣x) p(y∣x,z).p(y \mid x) = \sum_z p(z \mid x)\, p(y \mid x, z).

What it does. One generated response samples (or greedily decodes) a single zz, then an answer conditioned on it. It does not compute the sum. If the model writes a bad chain, the answer conditions on the bad chain. That is why sampling several chains and voting, self-consistency, approximates the sum better than one chain does.

Costs and failure modes

  • Tokens and latency. Every intermediate token is generated sequentially and billed. A chain of 2,000 tokens takes longer than a 20-token answer, whatever the hardware.
  • Error propagation. A wrong step early in the chain becomes context for everything after it.
  • Padding. Text such as “let me think carefully” adds tokens without adding a partial result. The reward lab shows what happens if length itself is rewarded.
  • Parsing. The final answer must be extracted from free text, which is a source of evaluation errors.
  • Faithfulness. Turpin and colleagues showed that biasing features in a prompt can change a model's answer without appearing in its chain of thought. Established See reasoning faithfulness.

Mini experiment

Ask any chat model a two-step word problem twice: once with “Answer with only the number” and once with “Show each calculation, then give the number”. Repeat on five problems. Count correct answers and output tokens for each mode. You have measured, on a tiny sample, the trade this page describes: tokens for accuracy. Treat five problems as an anecdote, not a benchmark.

Why should I care?

As a researcher

Nearly every reasoning method in the literature operates on chains of thought: they are sampled, voted on, searched over, scored step by step and rewarded.

As an engineer

Asking for intermediate work is the cheapest reasoning upgrade available, and it changes latency, token cost and how you must parse the final answer.

Modern systems that depend on it

  • self-consistency
  • tree search over thoughts
  • process reward models
  • reasoning models trained with RL

Historical context

Before

Prompts asked for the answer directly, and models that failed multi-step problems seemed to lack the capability altogether.

After

Models write intermediate steps before answering, by prompting at first and later by training, and the steps themselves became an object to sample, check and reward.

Used today

Reasoning models are trained to produce long intermediate work before answering; some products show it, others show only a summary.

What to remember

  • Intermediate text puts partial results into the context, where later tokens can use them.
  • Wei et al. (2022) used worked examples; Kojima et al. (2022) got large gains from “Let's think step by step” alone.
  • Prompting, fine-tuning on worked solutions and RL are three different ways to get a chain of thought; only the last two change weights.
  • Longer is not better: a chain can carry an early mistake forward, or pad without computing anything.
  • A plausible chain is not proof of how the answer was produced; check the result independently.

Key papers

Important

Large Language Models are Zero-Shot Reasoners

Takeshi Kojima, Shixiang Shane Gu et al. · 2022 · NeurIPS 2022

Showed that chain-of-thought behaviour needs no worked examples: one generic instruction elicits it, which made step-by-step prompting a default habit.

How to read it: The two-stage prompt (first elicit reasoning, then extract the answer) is a detail worth noticing: the answer still has to be parsed from free text.

~25 min readarXiv:2205.11916✓ verified 2026-10-06

Watch

3 h 31 min

Andrej Karpathy

Deep Dive into LLMs like ChatGPT

A long, general-audience walk through the whole pipeline behind a chat model, from internet text to tokens to pretraining to post-training.

Should know