Skip to content
Road to Intelligence

Part IV · Systems

Chapter 14

Reasoning Models

A little room to work. A better way to check.

1 h 30 min core path11 concepts3 interactivesCore path · 4 optional concepts foldedDeep · all 11 concepts shown in full

In one sentenceReasoning systems spend computation on intermediate work, alternatives and verification, and can be trained to make that work more useful.

The problem

A little room to work

Start at 1. You may add 2, multiply by 2, or subtract 3. Reach 19 in at most six moves.

The rules fit in one sentence. Finding a route is another matter. Always stepping to whatever is closest to 19 feels sensible: 1 → 3 → 6 → 12 → 14 → 16 → 18. Six moves, no solution: from 6 onward every number is even, an even number becomes odd only by subtracting 3, and subtracting always looks like a step backwards. The locally attractive choice at 3 discarded a useful alternative: 1 → 3 → 5 → 10 → 20 → 17 → 19.

This small puzzle exposes a distinction we will use throughout the chapter: knowing the available operations is different from allocating enough work to combine them. You can write intermediate results, explore alternatives, check candidates, or learn which attempts are worth making.

Chapter 13 spent computation interacting with an environment. Here we spend it before settling on an answer. A reasoning model can also use tools and act as an agent; the two ideas overlap. But more tool calls, more intermediate tokens, and a better-trained model are three different interventions.

A useful engineering definition of a reasoning model is a model trained to make productive use of intermediate computation on multi-step tasks. It is a description of training and behavior, not a claim of consciousness or a separate species of neural network. Interpretation

Two places to spend compute

Recall pretraining and post-training. Training changes parameters using many examples. Inference uses the resulting parameters to answer a particular request.

Extra inference computation does not, by itself, update the weights. A longer response runs the model more times; it does not automatically teach the checkpoint a lasting lesson. Established
Spend the computationWhat changesWhat survives the request?
Train on more or better dataThe model parametersA new checkpoint
Train with rewarded attemptsThe distribution of responsesA new checkpoint
Generate intermediate workThe current contextOnly what the application saves
Sample several solutionsThe candidate setOnly the chosen result or saved traces
Search and verifyWhich states or answers surviveOnly the recorded result and evidence

An ordinary autoregressive answer already takes repeated forward passes—roughly one per decoded token, ignoring serving optimizations. Reasoning does not turn “one forward pass” into “many” for the first time. It changes what those additional passes are used for. A brief answer may conceal a longer computation; a long answer may contain little useful work.

For a data engineer, think of training as investing in the system that generates query plans, and inference as the work spent on a particular query. Better planning and more execution effort can help together. Neither removes the need for a correct result.

The first move: intermediate work

Give the answer a scratchpad

Suppose six trays hold seven seedlings each, and half are moved outside. “21 remain” is the answer. “6 × 7 = 42; 42 ÷ 2 = 21” gives the next prediction useful intermediate state. The number 42 is now in the context instead of having to remain implicit in internal activations.

Chain-of-thought prompting demonstrated that examples containing intermediate steps improved performance on several reasoning tasks for sufficiently large models in the paper's experiments. Established A few months later, Kojima and colleagues found that adding only “Let's think step by step” before the answer, with no worked examples, raised one large model's accuracy on the MultiArith benchmark from 17.7% to 78.7%. Established Prompting a pretrained model with worked examples changes its input. Fine-tuning it on worked solutions changes its weights. RL changes those weights using a reward signal. Keep these three mechanisms separate even when they produce similar-looking text.

The probabilistic picture is the same autoregressive model from Chapter 8. Introduce intermediate text zz between question xx and answer yy:

p(y∣x)=∑zp(z∣x) p(y∣x,z).p(y\mid x)=\sum_z p(z\mid x)\,p(y\mid x,z).

This is a conceptual marginalization over possible intermediate sequences, not an algorithm that enumerates them all. One generated solution samples or decodes one such path. More tokens give additional computational steps and context, but also more chances to carry an early error forward.

Try the puzzle mentally again. The useful intermediate state is not “I should think carefully.” It is the current number, moves remaining, and alternatives you have not ruled out. Productive work changes what the next step can do.

Explore before committing

Chapter 1 introduced search over states. The same structure returns here: propose a continuation, evaluate it, retain promising possibilities and expand again. A beam keeps several states at each depth. A greedy search keeps only one.

Tree of Thoughts explored explicit search over intermediate text states, with evaluation and backtracking, on a small set of problem-solving tasks. Established That is an inference procedure around a model. It does not mean every model sold as a reasoning model runs an explicit tree search.

The lab uses arithmetic rather than generated text so every transition is inspectable. Press Play with one path. Then keep three paths. The extra branches consume checks; the successful route must still execute all six legal operations.

Try it · toy model

Search Before You Answer

Grow a real arithmetic search tree. Change its width and move budget, inspect discarded states, and discover why the nearest-looking path can fail.

Know well8 min

Our distance heuristic prefers 6 over 5 after reaching 3. It cannot see that 5 leads to 10, then 20, then 17 and 19. More width preserves that possibility. Reduce the move budget and width cannot compensate for missing depth. Try other targets: a successful demonstration on 19 is not evidence that three paths are optimal everywhere.

Deep diveWhere MCTS differsFrontier

Monte Carlo Tree Search repeatedly selects a path, expands a state, obtains an evaluation (often through a rollout or value estimate), and backs that value up through its ancestors. It balances exploiting high-value paths with exploring less-visited ones. Beam search instead retains a bounded frontier by a score. Our lab is beam search; it has no visit counts, rollouts or MCTS backup. For language-model search, defining a state, a useful evaluator and a stopping rule remains consequential design work.

Several answers, one decision

A crowd can repeat one mistake

Another way to spend compute is to start over. Generate several full solutions, extract their final answers, and choose the most frequent answer. Different wordings that end in 21 count toward the same answer.

Self-consistency improved results on the paper's arithmetic and commonsense benchmarks by sampling multiple reasoning paths and aggregating their answers. Established Agreement is useful when correct answers accumulate more reliably than any particular wrong answer. It is not a truth test.

Imagine eight runs that all overlook “half moved outside” and answer 42. Eight votes can make one misunderstanding look certain. Even separately sampled completions share a model, prompt and training history, so they can share a systematic error without literally copying text.

If each independent candidate solves a fixed task with probability pp, the chance that at least one succeeds is:

P(at least one correct)=1−(1−p)N.P(\text{at least one correct})=1-(1-p)^N.

At p=0.25p=0.25, four samples give 1−0.754≈68.4%1-0.75^4\approx68.4\%. That is availability, not the probability that your selection rule returns the correct answer. If no trustworthy selector identifies the good candidate, it may be wasted. The equation also assumes independent, identical trials on this task; plugging an average accuracy across mixed-difficulty tasks into it generally gives the wrong aggregate result.

Who checks the checker

A verifier evaluates a candidate. A calculator can check arithmetic. Tests can check a program on specified cases. A learned reward model can score a solution. These are different kinds of evidence with different blind spots.

Best-of-N generates N candidates and chooses the one with the highest verifier score. Self-consistency chooses by answer frequency. Cobbe and colleagues introduced the GSM8K grade-school math benchmark and trained verifiers to judge model solutions; generating many candidates and keeping the one the verifier ranked highest significantly improved performance on it. Established An oracle that simply tells you which candidates are truly correct is a useful experimental upper bound on selection from that batch, not a tool most applications possess.

In this lab, the correct answer is deliberately known. Each tile shows a candidate's answer and its verifier score, and a star marks the verifier's pick. You can see both what each selector chooses and whether the batch contained a better option. Press Play at the starting seed and watch the commonest mistake win the vote. Then make the verifier less reliable or the candidates more repetitive.

Try it · toy model

Many Answers, One Decision

Sample a batch, watch answers accumulate, and compare a vote with a noisy verifier. Correlate the attempts to see agreement stop being evidence.

Know well8 min

Separate a generation failure (“none of the candidates was correct”) from a selection failure (“a correct answer was present but discarded”). They suggest different fixes. Interpretation More samples address the first only when the generator has a chance of success. Improving the verifier may address the second.

The lab's verifier errs independently on each candidate, so more candidates mostly help it. Real scorers also have systematic blind spots, and then more candidates give them more chances to find a wrong answer they over-rate. Gao and colleagues measured this with a proxy reward model and a held-out “gold” one: pushing best-of-N harder against the proxy eventually lowered the gold score. Established

A test suite is stronger evidence than confident prose, but it checks only the cases and specification it encodes. A model can exploit a weak checker. For open-ended explanation, research design or ambiguous real-world decisions, defining a reward is itself part of the problem.

Reward the destination or the route

Return to 6×7÷26\times7\div2. Both of these responses end in 21:

  • Valid: 6 × 7 = 42; 42 ÷ 2 = 21.
  • Lucky: 6 × 7 = 40; 40 ÷ 2 = 21.

An outcome checker that looks only at the final number rewards both. A step checker rejects the second. Outcome supervision supplies a label for the result; process supervision supplies labels or scores for intermediate steps.

Let's Verify Step by Step found process supervision more effective than outcome supervision in its MATH evaluation and released step-level feedback data. Established This is evidence from a particular setup, not proof that process supervision always wins. Step labels can be expensive and ambiguous, and a learned process checker can be wrong.

More detail is not necessarily a better process. A verbose response can repeat an arithmetic mistake four times. Rewarding length, self-confidence or the phrase “let me verify” rewards observable surface features; none is equivalent to checking the arithmetic.

OptionalProcess vs Outcome Supervision· folded on the core path. Open it, or switch to Deep to show it here.

Teach the generator

Practice with a checkable reward

At inference time, we can choose among attempts. At training time, we can change the probability of those attempts. The reinforcement-learning loop is familiar from Chapter 5: sample behavior, obtain reward, update the policy. Here the behavior is a token sequence.

Reinforcement learning with verifiable rewards uses an executable or rule-based check to produce training feedback—for example, agreement with a known mathematical answer or passing specified code tests. Established “Verifiable” describes the reward mechanism. It does not guarantee that the specification is complete, the checker cannot be exploited, or the model will generalize.

The following lab actually updates four policy logits. It cannot invent a fifth response, so it isolates one question: which existing behavior does this reward make more likely? Train with each reward and compare correct-answer probability to valid-derivation probability.

Try it · toy model

You Get What You Reward

Train a four-response policy with real gradient updates. Reward outcomes, valid steps or length, and watch what becomes more likely.

Know well8 min
Deep diveThe exact gradient in the labShould know

Let θi\theta_i be the logit of response ii, pi=softmax⁡(θ)ip_i=\operatorname{softmax}(\theta)_i and rir_i its reward. The expected reward is J=∑ipiriJ=\sum_i p_i r_i. Its gradient is

∂J∂θi=pi(ri−J).\frac{\partial J}{\partial\theta_i}=p_i(r_i-J).

The lab applies θi←θi+ηpi(ri−J)\theta_i\leftarrow\theta_i+\eta p_i(r_i-J) with learning rate η=1\eta=1. Initially all four responses have probability 0.25. Under outcome reward, r=[1,1,0,0]r=[1,1,0,0] and J=0.5J=0.5, so the first two logits each increase by 0.125 and the others decrease by 0.125. The next distribution favors correct final answers while remaining unable to distinguish valid work from lucky work. This is exact expected-reward optimization in a tiny bandit; real LLM RL estimates gradients from sampled token sequences and often constrains updates.

Relative rewards and the training recipe

A reward of 1 is informative relative to something. If every sampled solution is correct—or every one is wrong—comparing that group supplies no within-group preference. A useful training problem must connect the model's current capabilities to a learnable signal.

DeepSeekMath introduced Group Relative Policy Optimization (GRPO), which uses group-relative rewards and avoids the separate critic used in common PPO implementations. Established Its group-normalized advantage can be written schematically as

Ai=ri−mean⁡(r)std⁡(r)+ϵ.A_i=\frac{r_i-\operatorname{mean}(r)}{\operatorname{std}(r)+\epsilon}.

This is the advantage signal, not the full GRPO objective. Probability ratios, clipping, token aggregation and regularization also matter. Implementations differ; a zero-variance group needs explicit handling. For rewards [1,0,0,1][1,0,0,1], the population mean and standard deviation are 0.5, giving advantages close to [1,−1,−1,1][1,-1,-1,1] when ϵ\epsilon is tiny.

The January 2025 DeepSeek-R1 report (its first arXiv version) distinguished R1-Zero, which applied RL without preliminary SFT, from R1, which used a multi-stage recipe including cold-start data. It also described transferring generated training data to smaller models. Established These are versioned historical examples, checked October 6, 2026—not a description of every current system. “RL alone” in this context still starts from a pretrained model; it does not mean learning language and mathematics from nothing.

Put together, the pieces of this chapter form three loops that run at different times and change different things. Step through each one.

Three loops: train, distill, serve

Changes the model's weights. Runs offline, before a checkpoint ships.

A teaching map, not any one lab’s pipeline. Real recipes interleave stages: DeepSeek-R1, for example, alternated supervised fine-tuning and RL, and its distilled models were trained on data the RL model generated.

How much an RL recipe adds new generalizable problem-solving behavior, versus amplifying behavior already available under sampling, depends on the model, data, task and evaluation budget. Active research A larger benchmark score alone cannot resolve that distinction. Inspect the base model's candidate distribution, matched-compute baselines and held-out task families.

OptionalGroup Relative Policy Optimization (GRPO)· folded on the core path. Open it, or switch to Deep to show it here.

Turn successful attempts into data

A slow, expensive generator can produce many attempts. Check them, retain useful solutions, and fine-tune another model on the retained question–solution pairs. This creates synthetic reasoning data. Training a smaller student on those traces is one form of reasoning distillation; it need not match the teacher's token probabilities.

The teacher can even be the student itself. STaR ran a simple loop: generate rationales for many questions, retry failures with the correct answer given as a hint, fine-tune on the rationales that ended in correct answers, and repeat. It improved over a model fine-tuned to predict answers directly on the paper's datasets. Established

Supervised training on selected teacher outputs changes the student's weights; presenting the same outputs as prompt examples does not. Established The teacher need not be a single model call. It can be a model plus search, tools and filtering. The training record should say which system produced the targets.

The data-pipeline analogy is direct: generate → validate → filter → deduplicate → split → train → evaluate. Keep held-out questions and their near-duplicates out of both generation prompts and student training. A correct final answer attached to a flawed rationale can still be a bad teaching example. Filtering only by final answer inherits the checker problem from the previous lab.

Distillation moves some expense upstream: spend more generating training targets so the student may need less work on later requests. It does not guarantee the student's competence, faithfulness or efficiency on new tasks. Compare it with the unmodified student using the same evaluation and inference budget.

OptionalSynthetic Reasoning Data and Distillation· folded on the core path. Open it, or switch to Deep to show it here.

From prompt trick to product

Every mechanism in this chapter appeared in research papers before anyone sold a “reasoning model”. Worked examples and verifiers came first, then voting, search and step-level rewards, then group-relative RL.

The road to reasoning models, 2021–2025Open the full timeline →

OpenAI's September 2024 announcement of o1 described a model trained with large-scale reinforcement learning to use a chain of thought, and reported that its performance improved with both more RL training compute and more time spent thinking at test time. Established The announcement gave few training details; checked October 6, 2026.

Its AIME 2024 results put this chapter's selection ideas side by side. OpenAI reported that o1 solved 74% of problems with one sample per problem, 83% with a vote over 64 samples, and 93% when re-ranking 1,000 samples with a learned scoring function. Established One generator, three selection rules, three numbers, and the inference cost rises at each step. These are the company's own evaluations, not an independent measurement.

The same announcement said users would not see the raw chain of thought, only a model-generated summary of it. The text a user reads is then not even the intermediate sequence the model produced, which matters for the last section of this chapter. Four months later, DeepSeek-R1 published open weights and a written training account, so the recipe could be studied outside one company.

“Reasoning model” became a product category with these releases, but the ingredients are older: the 2024–2025 change was mainly training intermediate work with checkable rewards at scale and budgeting inference for it. Interpretation

Spend carefully

More compute needs a useful place to go

There are at least three inference budgets: depth within a solution, breadth across solutions, and verification to select or revise them. A token cap controls only part of that system. Parallel candidates can reduce wall-clock delay while consuming more total compute; a long sequential chain cannot gain that same parallel speedup.

Snell and colleagues found that effective test-time compute allocation depended on problem difficulty and the inference strategy in their experiments. Established Their result motivates choosing the allocation by task; it does not establish a universal curve or make inference compute interchangeable with training compute.

Our search lab shows a narrow version of this: width saves a discarded path, depth permits enough operations, and an exact check recognizes the target. Giving the wrong heuristic more budget can still disappoint. Likewise, a language model can spend tokens repeating a misconception.

For a rough planning ledger, if a model costs CtrainC_{\text{train}} to train and serves QQ requests at average inference cost CinferC_{\text{infer}}, then

Ctotal=Ctrain+Q Cinfer.C_{\text{total}}=C_{\text{train}}+Q\,C_{\text{infer}}.

The units must match—FLOPs, energy or money under a specified costing model. This is accounting, not a performance law. Quality, latency, memory and request difficulty constrain which allocations are useful. Spending 10 extra units per request adds 10 million units across a million requests; an upstream improvement can therefore be worth amortizing if it actually saves work at matched quality.

Evaluate a budget policy, not just a model name: which requests receive more work, when does the system stop, and what does it do when the budget runs out? Interpretation An answer that states what remains unresolved is preferable to treating exhausted tokens as evidence of success.

Read the evidence rather than the monologue

A long worked answer is an artifact you can inspect. It is not a guaranteed explanation of the computation that caused the answer. A system may expose a summary, hide intermediate tokens, or generate a plausible rationale after being influenced by something it never mentions.

Turpin and colleagues demonstrated settings where biasing information changed model answers without appearing in the chain-of-thought explanation. Established The conclusion is specific: a convincing explanation alone does not establish faithfulness. It does not imply that all intermediate reasoning is useless or always misleading.

OptionalFaithfulness of Reasoning Traces· folded on the core path. Open it, or switch to Deep to show it here.

Before trusting a reasoning result, make five questions concrete:

  1. What changed? Prompt, weights, search, tools, verifier, or all of them?
  2. What was the budget? Count generated tokens, repeated prompt processing, verifier work and latency.
  3. What was selected? First answer, vote, best scored answer, or an oracle-picked success?
  4. What was held out? Check contamination, task diversity and uncertainty across runs.
  5. What was actually verified? A final number, intermediate steps, tests, or an open-ended human judgment?
Reasoning systems are useful when extra work buys better outcomes under a well-specified check and budget. The amount of text produced is not that outcome. Interpretation

You have now connected search from Chapter 1, autoregressive generation from Chapter 8, policy optimization from Chapter 10, serving costs from Chapter 11, and verification from Chapter 13. Chapter 15 extends the inputs and outputs beyond text. The same questions remain when the model reasons over an image, a recording or a physical action: what did it observe, what did it compute, and what evidence supports the answer?

Concepts in this chapter

Mark each one as you go. Must-know concepts are the core path.

What do I actually need to remember?

  • Training changes weights; ordinary inference changes the current context and candidate set.
  • A normal autoregressive answer already uses repeated forward passes.
  • Intermediate work can help later predictions; longer text is not inherently better reasoning.
  • Search spends depth and breadth to preserve alternatives, using a fallible evaluator.
  • Self-consistency aggregates answers; shared errors can make agreement misleading.
  • At least one correct candidate is not the same as selecting a correct answer.
  • Outcome rewards judge results; process supervision supplies intermediate feedback.
  • Verifiable rewards are only as complete and robust as their checkers.
  • Distillation trains on selected teacher outputs, moving some expense upstream.
  • Compare quality, cost and latency; a convincing trace does not guarantee faithfulness.

You do not need to memorize everything else. This list is the revision sheet.

Key papers

Essential

Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Jason Wei, Xuezhi Wang et al. · 2022 · NeurIPS 2022

Showed that prompting large models to write out intermediate steps markedly improves multi-step reasoning — the seed of today's reasoning models.

Problem
Large models often failed at arithmetic and multi-step reasoning when asked for the answer directly.
What was new
Few-shot examples that include step-by-step reasoning, which large enough models imitate.
~40 min readarXiv:2201.11903✓ verified 2026-09-26
Important

Large Language Models are Zero-Shot Reasoners

Takeshi Kojima, Shixiang Shane Gu et al. · 2022 · NeurIPS 2022

Showed that chain-of-thought behaviour needs no worked examples: one generic instruction elicits it, which made step-by-step prompting a default habit.

Problem
Chain-of-thought prompting seemed to need hand-written, task-specific worked examples.
What was new
Appending “Let's think step by step” before the answer raised MultiArith accuracy from 17.7% to 78.7% and GSM8K from 10.4% to 40.7% for InstructGPT (text-davinci-002), with similar gains for PaLM 540B.

How to read it: The two-stage prompt (first elicit reasoning, then extract the answer) is a detail worth noticing: the answer still has to be parsed from free text.

~25 min readarXiv:2205.11916✓ verified 2026-10-06
Important

Training Verifiers to Solve Math Word Problems

Karl Cobbe, Vineet Kosaraju et al. · 2021

Introduced GSM8K, the grade-school math benchmark used for years afterwards, and the generate-many-then-verify recipe that best-of-N selection still follows.

Problem
Even the largest language models of the time failed at multi-step grade-school arithmetic, and one sampled solution was often wrong.
What was new
A dataset of 8.5K word problems, and a trained verifier that scores sampled solutions so the highest-ranked one is kept; verification scaled better with more data than fine-tuning alone.

How to read it: Read the sections on the verifier and on how test performance changes with the number of sampled completions; that trade-off is Chapter 14's selection problem.

~35 min readarXiv:2110.14168✓ verified 2026-10-06
Essential

Self-Consistency Improves Chain of Thought Reasoning in Language Models

Xuezhi Wang, Jason Wei et al. · 2022 · ICLR 2023

Turned extra inference compute into accuracy without any training: sample many reasoning paths and keep the answer they most often reach.

Problem
Greedy decoding commits to one reasoning path, and a single early mistake decides the answer.
What was new
Replace greedy chain-of-thought decoding with sampling plus a vote over final answers; reported gains included +17.9% on GSM8K and +11.0% on SVAMP.

How to read it: Look at how accuracy grows with the number of sampled paths, and notice the method needs answers that can be compared exactly.

~35 min readarXiv:2203.11171✓ verified 2026-10-06
Important

Tree of Thoughts: Deliberate Problem Solving with Large Language Models

Shunyu Yao, Dian Yu et al. · 2023 · NeurIPS 2023

Made classical search explicit around a language model: propose partial solutions, evaluate them, and backtrack instead of committing left to right.

Problem
Token-by-token, left-to-right generation cannot look ahead or undo an early choice on tasks that need exploration.
What was new
Search (breadth- or depth-first) over “thoughts”, coherent chunks of text, with the model also evaluating states; on Game of 24, GPT-4 went from 4% with chain of thought to 74%.

How to read it: For each task, write down what a state is, who scores it and how many model calls a solution costs; the gains come with a large call budget.

~40 min readarXiv:2305.10601✓ verified 2026-10-06
Important

Let's Verify Step by Step

Hunter Lightman, Vineet Kosaraju et al. · 2023

The clearest comparison of rewarding steps versus rewarding final answers, and the source of PRM800K, a widely used set of step-level labels.

Problem
A final-answer label cannot say where a derivation went wrong, and it rewards lucky answers reached by invalid steps.
What was new
A process-supervised reward model, trained on 800,000 human step labels, selected solutions better than an outcome-supervised one on MATH; the best solved 78% of a representative test subset.

How to read it: Both reward models are used to pick the best of many samples (best-of-N), not to train the generator with RL; keep that setup in mind when generalising.

~45 min readarXiv:2305.20050✓ verified 2026-10-06
Important

Scaling Laws for Reward Model Overoptimization

Leo Gao, John Schulman, Jacob Hilton · 2022

Explains why maximizing a learned judge can diverge from the intended outcome.

Problem
A policy can exploit errors in a proxy reward model.
What was new
Study optimization against a proxy while measuring a synthetic gold reward.

How to read it: Read the synthetic-gold-reward setup before interpreting the scaling curves.

~30 min readarXiv:2210.10760✓ verified 2026-10-04
Important

DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Zhihong Shao, Peiyi Wang et al. · 2024

Introduced GRPO, the group-relative RL method later used to train DeepSeek-R1 and widely adopted for RL with verifiable rewards.

Problem
PPO needs a separate value model as large as the policy, which makes RL on long responses memory-hungry.
What was new
A 7B model further pre-trained on 120B math tokens, plus GRPO: a PPO variant that baselines each response against a group of responses to the same prompt instead of a learned critic.

How to read it: For GRPO, go to the RL section and compare its objective with PPO's term by term; the data pipeline sections are a separate, also useful, story.

~1 h readarXiv:2402.03300✓ verified 2026-10-06
Important

STaR: Bootstrapping Reasoning With Reasoning

Eric Zelikman, Yuhuai Wu et al. · 2022 · NeurIPS 2022

An early, clear version of the loop behind synthetic reasoning data: a model generates rationales, keeps the ones that reach correct answers, and trains on them.

Problem
Training a model to write rationales seemed to need large human-written rationale datasets.
What was new
The Self-Taught Reasoner loop: generate rationales, retry failures with the correct answer as a hint (“rationalization”), fine-tune on rationales that ended correctly, repeat.

How to read it: Notice the filter is the final answer only, so a rationale that reaches the right answer by a flawed route is kept; compare Chapter 14's lucky answer.

~35 min readarXiv:2203.14465✓ verified 2026-10-06
Essential

Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Charlie Snell, Jaehoon Lee et al. · 2024

Treated inference compute as a budget to allocate, and showed the right allocation depends on how hard the prompt is.

Problem
Earlier results on spending more inference compute were mixed, with little guidance on which method to use when.
What was new
Compared searching against process-based verifiers with revising answers sequentially; allocating compute per prompt by difficulty was over 4× more efficient than best-of-N, and in a FLOPs-matched comparison could beat a 14× larger model on problems the small model sometimes solved.

How to read it: Keep the condition attached to the headline: test-time compute beat a larger model only where the smaller model already had non-trivial success.

~1 h readarXiv:2408.03314✓ verified 2026-10-06
Important

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

DeepSeek-AI et al. · 2025

An openly released reasoning model, with a detailed account of training long chains of reasoning mainly through reinforcement learning on verifiable problems.

Problem
How reasoning models were trained was largely undisclosed.
What was new
Large-scale RL with rule-based rewards on math and code, plus distillation of the resulting reasoning into smaller open models.

How to read it: Read the R1-Zero and R1 sections separately: the first is RL straight from a base model, the second a multi-stage recipe with cold-start data, and the distillation results are a third story.

~1 h readarXiv:2501.12948✓ verified 2026-09-26
Important

Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting

Miles Turpin, Julian Michael et al. · 2023 · NeurIPS 2023

Showed with controlled experiments that a chain of thought can omit what actually drove the answer, so a plausible explanation is not evidence of faithfulness.

Problem
It is tempting to read a model's step-by-step explanation as the reason for its answer.
What was new
Biasing features, such as always making the few-shot answer “(A)”, changed answers without being mentioned; explanations rationalised the biased answers, and accuracy fell by up to 36% on 13 BIG-Bench Hard tasks.

How to read it: The experimental design is the lesson: change an input the explanation never mentions and see whether the answer moves.

~35 min readarXiv:2305.04388✓ verified 2026-10-06

Watch

3 h 31 min

Andrej Karpathy

Deep Dive into LLMs like ChatGPT

A long, general-audience walk through the whole pipeline behind a chat model, from internet text to tokens to pretraining to post-training.

Covers: The full training stack behind a chat model, from internet text and tokens to a pretrained base model and the post-training that turns it into an assistant.

Should know

What came next?

Chapter 15

Multimodal AI

The world isn't made of text.

This chapter is being written.