Part IV · Systems
Chapter 14
Reasoning Models
A little room to work. A better way to check.
In one sentenceReasoning systems spend computation on intermediate work, alternatives and verification, and can be trained to make that work more useful.
The problem
A little room to work
Start at 1. You may add 2, multiply by 2, or subtract 3. Reach 19 in at most six moves.
The rules fit in one sentence. Finding a route is another matter. Always stepping to whatever is closest to 19 feels sensible: 1 → 3 → 6 → 12 → 14 → 16 → 18. Six moves, no solution: from 6 onward every number is even, an even number becomes odd only by subtracting 3, and subtracting always looks like a step backwards. The locally attractive choice at 3 discarded a useful alternative: 1 → 3 → 5 → 10 → 20 → 17 → 19.
This small puzzle exposes a distinction we will use throughout the chapter: knowing the available operations is different from allocating enough work to combine them. You can write intermediate results, explore alternatives, check candidates, or learn which attempts are worth making.
Chapter 13 spent computation interacting with an environment. Here we spend it before settling on an answer. A reasoning model can also use tools and act as an agent; the two ideas overlap. But more tool calls, more intermediate tokens, and a better-trained model are three different interventions.
A useful engineering definition of a reasoning model is a model trained to make productive use of intermediate computation on multi-step tasks. It is a description of training and behavior, not a claim of consciousness or a separate species of neural network. InterpretationTwo places to spend compute
Recall pretraining and post-training. Training changes parameters using many examples. Inference uses the resulting parameters to answer a particular request.
Extra inference computation does not, by itself, update the weights. A longer response runs the model more times; it does not automatically teach the checkpoint a lasting lesson. Established| Spend the computation | What changes | What survives the request? |
|---|---|---|
| Train on more or better data | The model parameters | A new checkpoint |
| Train with rewarded attempts | The distribution of responses | A new checkpoint |
| Generate intermediate work | The current context | Only what the application saves |
| Sample several solutions | The candidate set | Only the chosen result or saved traces |
| Search and verify | Which states or answers survive | Only the recorded result and evidence |
An ordinary autoregressive answer already takes repeated forward passes—roughly one per decoded token, ignoring serving optimizations. Reasoning does not turn “one forward pass” into “many” for the first time. It changes what those additional passes are used for. A brief answer may conceal a longer computation; a long answer may contain little useful work.
For a data engineer, think of training as investing in the system that generates query plans, and inference as the work spent on a particular query. Better planning and more execution effort can help together. Neither removes the need for a correct result.
The first move: intermediate work
Give the answer a scratchpad
Suppose six trays hold seven seedlings each, and half are moved outside. “21 remain” is the answer. “6 × 7 = 42; 42 ÷ 2 = 21” gives the next prediction useful intermediate state. The number 42 is now in the context instead of having to remain implicit in internal activations.
Chain-of-thought prompting demonstrated that examples containing intermediate steps improved performance on several reasoning tasks for sufficiently large models in the paper's experiments. Established A few months later, Kojima and colleagues found that adding only “Let's think step by step” before the answer, with no worked examples, raised one large model's accuracy on the MultiArith benchmark from 17.7% to 78.7%. Established Prompting a pretrained model with worked examples changes its input. Fine-tuning it on worked solutions changes its weights. RL changes those weights using a reward signal. Keep these three mechanisms separate even when they produce similar-looking text.
The probabilistic picture is the same autoregressive model from Chapter 8. Introduce intermediate text between question and answer :
This is a conceptual marginalization over possible intermediate sequences, not an algorithm that enumerates them all. One generated solution samples or decodes one such path. More tokens give additional computational steps and context, but also more chances to carry an early error forward.
Try the puzzle mentally again. The useful intermediate state is not “I should think carefully.” It is the current number, moves remaining, and alternatives you have not ruled out. Productive work changes what the next step can do.
Explore before committing
Chapter 1 introduced search over states. The same structure returns here: propose a continuation, evaluate it, retain promising possibilities and expand again. A beam keeps several states at each depth. A greedy search keeps only one.
Tree of Thoughts explored explicit search over intermediate text states, with evaluation and backtracking, on a small set of problem-solving tasks. Established That is an inference procedure around a model. It does not mean every model sold as a reasoning model runs an explicit tree search.
The lab uses arithmetic rather than generated text so every transition is inspectable. Press Play with one path. Then keep three paths. The extra branches consume checks; the successful route must still execute all six legal operations.
Try it · toy model
Grow a real arithmetic search tree. Change its width and move budget, inspect discarded states, and discover why the nearest-looking path can fail.
Our distance heuristic prefers 6 over 5 after reaching 3. It cannot see that 5 leads to 10, then 20, then 17 and 19. More width preserves that possibility. Reduce the move budget and width cannot compensate for missing depth. Try other targets: a successful demonstration on 19 is not evidence that three paths are optimal everywhere.
Deep diveWhere MCTS differsFrontier
Monte Carlo Tree Search repeatedly selects a path, expands a state, obtains an evaluation (often through a rollout or value estimate), and backs that value up through its ancestors. It balances exploiting high-value paths with exploring less-visited ones. Beam search instead retains a bounded frontier by a score. Our lab is beam search; it has no visit counts, rollouts or MCTS backup. For language-model search, defining a state, a useful evaluator and a stopping rule remains consequential design work.
Several answers, one decision
A crowd can repeat one mistake
Another way to spend compute is to start over. Generate several full solutions, extract their final answers, and choose the most frequent answer. Different wordings that end in 21 count toward the same answer.
Self-consistency improved results on the paper's arithmetic and commonsense benchmarks by sampling multiple reasoning paths and aggregating their answers. Established Agreement is useful when correct answers accumulate more reliably than any particular wrong answer. It is not a truth test.
Imagine eight runs that all overlook “half moved outside” and answer 42. Eight votes can make one misunderstanding look certain. Even separately sampled completions share a model, prompt and training history, so they can share a systematic error without literally copying text.
If each independent candidate solves a fixed task with probability , the chance that at least one succeeds is:
At , four samples give . That is availability, not the probability that your selection rule returns the correct answer. If no trustworthy selector identifies the good candidate, it may be wasted. The equation also assumes independent, identical trials on this task; plugging an average accuracy across mixed-difficulty tasks into it generally gives the wrong aggregate result.
Who checks the checker
A verifier evaluates a candidate. A calculator can check arithmetic. Tests can check a program on specified cases. A learned reward model can score a solution. These are different kinds of evidence with different blind spots.
Best-of-N generates N candidates and chooses the one with the highest verifier score. Self-consistency chooses by answer frequency. Cobbe and colleagues introduced the GSM8K grade-school math benchmark and trained verifiers to judge model solutions; generating many candidates and keeping the one the verifier ranked highest significantly improved performance on it. Established An oracle that simply tells you which candidates are truly correct is a useful experimental upper bound on selection from that batch, not a tool most applications possess.
In this lab, the correct answer is deliberately known. Each tile shows a candidate's answer and its verifier score, and a star marks the verifier's pick. You can see both what each selector chooses and whether the batch contained a better option. Press Play at the starting seed and watch the commonest mistake win the vote. Then make the verifier less reliable or the candidates more repetitive.
Try it · toy model
Sample a batch, watch answers accumulate, and compare a vote with a noisy verifier. Correlate the attempts to see agreement stop being evidence.
Separate a generation failure (“none of the candidates was correct”) from a selection failure (“a correct answer was present but discarded”). They suggest different fixes. Interpretation More samples address the first only when the generator has a chance of success. Improving the verifier may address the second.
The lab's verifier errs independently on each candidate, so more candidates mostly help it. Real scorers also have systematic blind spots, and then more candidates give them more chances to find a wrong answer they over-rate. Gao and colleagues measured this with a proxy reward model and a held-out “gold” one: pushing best-of-N harder against the proxy eventually lowered the gold score. Established
A test suite is stronger evidence than confident prose, but it checks only the cases and specification it encodes. A model can exploit a weak checker. For open-ended explanation, research design or ambiguous real-world decisions, defining a reward is itself part of the problem.
Reward the destination or the route
Return to . Both of these responses end in 21:
- Valid: 6 × 7 = 42; 42 ÷ 2 = 21.
- Lucky: 6 × 7 = 40; 40 ÷ 2 = 21.
An outcome checker that looks only at the final number rewards both. A step checker rejects the second. Outcome supervision supplies a label for the result; process supervision supplies labels or scores for intermediate steps.
Let's Verify Step by Step found process supervision more effective than outcome supervision in its MATH evaluation and released step-level feedback data. Established This is evidence from a particular setup, not proof that process supervision always wins. Step labels can be expensive and ambiguous, and a learned process checker can be wrong.
More detail is not necessarily a better process. A verbose response can repeat an arithmetic mistake four times. Rewarding length, self-confidence or the phrase “let me verify” rewards observable surface features; none is equivalent to checking the arithmetic.
Teach the generator
Practice with a checkable reward
At inference time, we can choose among attempts. At training time, we can change the probability of those attempts. The reinforcement-learning loop is familiar from Chapter 5: sample behavior, obtain reward, update the policy. Here the behavior is a token sequence.
Reinforcement learning with verifiable rewards uses an executable or rule-based check to produce training feedback—for example, agreement with a known mathematical answer or passing specified code tests. Established “Verifiable” describes the reward mechanism. It does not guarantee that the specification is complete, the checker cannot be exploited, or the model will generalize.
The following lab actually updates four policy logits. It cannot invent a fifth response, so it isolates one question: which existing behavior does this reward make more likely? Train with each reward and compare correct-answer probability to valid-derivation probability.
Try it · toy model
Train a four-response policy with real gradient updates. Reward outcomes, valid steps or length, and watch what becomes more likely.
Deep diveThe exact gradient in the labShould know
Let be the logit of response , and its reward. The expected reward is . Its gradient is
The lab applies with learning rate . Initially all four responses have probability 0.25. Under outcome reward, and , so the first two logits each increase by 0.125 and the others decrease by 0.125. The next distribution favors correct final answers while remaining unable to distinguish valid work from lucky work. This is exact expected-reward optimization in a tiny bandit; real LLM RL estimates gradients from sampled token sequences and often constrains updates.
Relative rewards and the training recipe
A reward of 1 is informative relative to something. If every sampled solution is correct—or every one is wrong—comparing that group supplies no within-group preference. A useful training problem must connect the model's current capabilities to a learnable signal.
DeepSeekMath introduced Group Relative Policy Optimization (GRPO), which uses group-relative rewards and avoids the separate critic used in common PPO implementations. Established Its group-normalized advantage can be written schematically as
This is the advantage signal, not the full GRPO objective. Probability ratios, clipping, token aggregation and regularization also matter. Implementations differ; a zero-variance group needs explicit handling. For rewards , the population mean and standard deviation are 0.5, giving advantages close to when is tiny.
The January 2025 DeepSeek-R1 report (its first arXiv version) distinguished R1-Zero, which applied RL without preliminary SFT, from R1, which used a multi-stage recipe including cold-start data. It also described transferring generated training data to smaller models. Established These are versioned historical examples, checked October 6, 2026—not a description of every current system. “RL alone” in this context still starts from a pretrained model; it does not mean learning language and mathematics from nothing.
Put together, the pieces of this chapter form three loops that run at different times and change different things. Step through each one.
Changes the model's weights. Runs offline, before a checkpoint ships.
A teaching map, not any one lab’s pipeline. Real recipes interleave stages: DeepSeek-R1, for example, alternated supervised fine-tuning and RL, and its distilled models were trained on data the RL model generated.
How much an RL recipe adds new generalizable problem-solving behavior, versus amplifying behavior already available under sampling, depends on the model, data, task and evaluation budget. Active research A larger benchmark score alone cannot resolve that distinction. Inspect the base model's candidate distribution, matched-compute baselines and held-out task families.
Turn successful attempts into data
A slow, expensive generator can produce many attempts. Check them, retain useful solutions, and fine-tune another model on the retained question–solution pairs. This creates synthetic reasoning data. Training a smaller student on those traces is one form of reasoning distillation; it need not match the teacher's token probabilities.
The teacher can even be the student itself. STaR ran a simple loop: generate rationales for many questions, retry failures with the correct answer given as a hint, fine-tune on the rationales that ended in correct answers, and repeat. It improved over a model fine-tuned to predict answers directly on the paper's datasets. Established
Supervised training on selected teacher outputs changes the student's weights; presenting the same outputs as prompt examples does not. Established The teacher need not be a single model call. It can be a model plus search, tools and filtering. The training record should say which system produced the targets.
The data-pipeline analogy is direct: generate → validate → filter → deduplicate → split → train → evaluate. Keep held-out questions and their near-duplicates out of both generation prompts and student training. A correct final answer attached to a flawed rationale can still be a bad teaching example. Filtering only by final answer inherits the checker problem from the previous lab.
Distillation moves some expense upstream: spend more generating training targets so the student may need less work on later requests. It does not guarantee the student's competence, faithfulness or efficiency on new tasks. Compare it with the unmodified student using the same evaluation and inference budget.
From prompt trick to product
Every mechanism in this chapter appeared in research papers before anyone sold a “reasoning model”. Worked examples and verifiers came first, then voting, search and step-level rewards, then group-relative RL.
OpenAI's September 2024 announcement of o1 described a model trained with large-scale reinforcement learning to use a chain of thought, and reported that its performance improved with both more RL training compute and more time spent thinking at test time. Established The announcement gave few training details; checked October 6, 2026.
Its AIME 2024 results put this chapter's selection ideas side by side. OpenAI reported that o1 solved 74% of problems with one sample per problem, 83% with a vote over 64 samples, and 93% when re-ranking 1,000 samples with a learned scoring function. Established One generator, three selection rules, three numbers, and the inference cost rises at each step. These are the company's own evaluations, not an independent measurement.
The same announcement said users would not see the raw chain of thought, only a model-generated summary of it. The text a user reads is then not even the intermediate sequence the model produced, which matters for the last section of this chapter. Four months later, DeepSeek-R1 published open weights and a written training account, so the recipe could be studied outside one company.
“Reasoning model” became a product category with these releases, but the ingredients are older: the 2024–2025 change was mainly training intermediate work with checkable rewards at scale and budgeting inference for it. InterpretationSpend carefully
More compute needs a useful place to go
There are at least three inference budgets: depth within a solution, breadth across solutions, and verification to select or revise them. A token cap controls only part of that system. Parallel candidates can reduce wall-clock delay while consuming more total compute; a long sequential chain cannot gain that same parallel speedup.
Snell and colleagues found that effective test-time compute allocation depended on problem difficulty and the inference strategy in their experiments. Established Their result motivates choosing the allocation by task; it does not establish a universal curve or make inference compute interchangeable with training compute.
Our search lab shows a narrow version of this: width saves a discarded path, depth permits enough operations, and an exact check recognizes the target. Giving the wrong heuristic more budget can still disappoint. Likewise, a language model can spend tokens repeating a misconception.
For a rough planning ledger, if a model costs to train and serves requests at average inference cost , then
The units must match—FLOPs, energy or money under a specified costing model. This is accounting, not a performance law. Quality, latency, memory and request difficulty constrain which allocations are useful. Spending 10 extra units per request adds 10 million units across a million requests; an upstream improvement can therefore be worth amortizing if it actually saves work at matched quality.
Evaluate a budget policy, not just a model name: which requests receive more work, when does the system stop, and what does it do when the budget runs out? Interpretation An answer that states what remains unresolved is preferable to treating exhausted tokens as evidence of success.
Read the evidence rather than the monologue
A long worked answer is an artifact you can inspect. It is not a guaranteed explanation of the computation that caused the answer. A system may expose a summary, hide intermediate tokens, or generate a plausible rationale after being influenced by something it never mentions.
Turpin and colleagues demonstrated settings where biasing information changed model answers without appearing in the chain-of-thought explanation. Established The conclusion is specific: a convincing explanation alone does not establish faithfulness. It does not imply that all intermediate reasoning is useless or always misleading.
Before trusting a reasoning result, make five questions concrete:
- What changed? Prompt, weights, search, tools, verifier, or all of them?
- What was the budget? Count generated tokens, repeated prompt processing, verifier work and latency.
- What was selected? First answer, vote, best scored answer, or an oracle-picked success?
- What was held out? Check contamination, task diversity and uncertainty across runs.
- What was actually verified? A final number, intermediate steps, tests, or an open-ended human judgment?
You have now connected search from Chapter 1, autoregressive generation from Chapter 8, policy optimization from Chapter 10, serving costs from Chapter 11, and verification from Chapter 13. Chapter 15 extends the inputs and outputs beyond text. The same questions remain when the model reasons over an image, a recording or a physical action: what did it observe, what did it compute, and what evidence supports the answer?
Concepts in this chapter
Mark each one as you go. Must-know concepts are the core path.
- Chain of ThoughtA chain of thought is intermediate text generated before the final answer, so that later tokens can condition on partial results instead of computing everything implicitly in one step.Know wellMust know
- Inference-Time ScalingInference-time scaling is improving answers by spending more computation per request, on longer reasoning, more samples, search or verification, and it pays off only when that computation is aimed by a useful selector and matched to the problem's difficulty.Know wellMust know
- Search over Reasoning StepsSearch over reasoning steps keeps several partial solutions alive, scores them, and expands the most promising, so that one tempting early choice does not decide the whole answer.Know wellMust know
- Verifiers and Best-of-NA verifier scores or checks candidate solutions so a system can keep the best one, and the quality of that check, not the number of candidates, often decides the final accuracy.Know wellMust know
- Self-ConsistencySelf-consistency samples several independent reasoning paths for the same question and returns the final answer they most often agree on.Know wellMust know
- Training vs Inference ComputeTraining compute changes a model's parameters once, for every future request; inference compute is spent on one request with the parameters held fixed, and nothing it learns survives unless someone saves it.Know wellMust know
- RL with Verifiable RewardsReinforcement learning with verifiable rewards trains a model on problems whose results a program can check, such as math answers or code tests, and raises the probability of the responses that pass.Know wellMust know
- Process vs Outcome SupervisionOutcome supervision gives feedback on a solution's final result; process supervision gives feedback on each intermediate step, so a lucky answer reached by invalid steps is not rewarded.Know wellShould know
- Synthetic Reasoning Data and DistillationReasoning distillation fine-tunes a model on solutions that a stronger or more expensive system generated and that passed a check, moving inference effort upstream into training data.Know wellShould know
- Faithfulness of Reasoning TracesA reasoning trace is faithful to the extent that it accurately reflects what actually caused the model's answer; a trace can be useful, correct and convincing while still leaving out what mattered.Know wellShould know
- Group Relative Policy Optimization (GRPO)GRPO is a PPO-style policy-optimization method that judges each sampled response against the other responses to the same prompt, replacing PPO's learned value model with a group average.Know wellFrontier
What do I actually need to remember?
- Training changes weights; ordinary inference changes the current context and candidate set.
- A normal autoregressive answer already uses repeated forward passes.
- Intermediate work can help later predictions; longer text is not inherently better reasoning.
- Search spends depth and breadth to preserve alternatives, using a fallible evaluator.
- Self-consistency aggregates answers; shared errors can make agreement misleading.
- At least one correct candidate is not the same as selecting a correct answer.
- Outcome rewards judge results; process supervision supplies intermediate feedback.
- Verifiable rewards are only as complete and robust as their checkers.
- Distillation trains on selected teacher outputs, moving some expense upstream.
- Compare quality, cost and latency; a convincing trace does not guarantee faithfulness.
You do not need to memorize everything else. This list is the revision sheet.
Key papers
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Jason Wei, Xuezhi Wang et al. · 2022 · NeurIPS 2022
Showed that prompting large models to write out intermediate steps markedly improves multi-step reasoning — the seed of today's reasoning models.
- Problem
- Large models often failed at arithmetic and multi-step reasoning when asked for the answer directly.
- What was new
- Few-shot examples that include step-by-step reasoning, which large enough models imitate.
Large Language Models are Zero-Shot Reasoners
Takeshi Kojima, Shixiang Shane Gu et al. · 2022 · NeurIPS 2022
Showed that chain-of-thought behaviour needs no worked examples: one generic instruction elicits it, which made step-by-step prompting a default habit.
- Problem
- Chain-of-thought prompting seemed to need hand-written, task-specific worked examples.
- What was new
- Appending “Let's think step by step” before the answer raised MultiArith accuracy from 17.7% to 78.7% and GSM8K from 10.4% to 40.7% for InstructGPT (text-davinci-002), with similar gains for PaLM 540B.
How to read it: The two-stage prompt (first elicit reasoning, then extract the answer) is a detail worth noticing: the answer still has to be parsed from free text.
Training Verifiers to Solve Math Word Problems
Karl Cobbe, Vineet Kosaraju et al. · 2021
Introduced GSM8K, the grade-school math benchmark used for years afterwards, and the generate-many-then-verify recipe that best-of-N selection still follows.
- Problem
- Even the largest language models of the time failed at multi-step grade-school arithmetic, and one sampled solution was often wrong.
- What was new
- A dataset of 8.5K word problems, and a trained verifier that scores sampled solutions so the highest-ranked one is kept; verification scaled better with more data than fine-tuning alone.
How to read it: Read the sections on the verifier and on how test performance changes with the number of sampled completions; that trade-off is Chapter 14's selection problem.
Self-Consistency Improves Chain of Thought Reasoning in Language Models
Xuezhi Wang, Jason Wei et al. · 2022 · ICLR 2023
Turned extra inference compute into accuracy without any training: sample many reasoning paths and keep the answer they most often reach.
- Problem
- Greedy decoding commits to one reasoning path, and a single early mistake decides the answer.
- What was new
- Replace greedy chain-of-thought decoding with sampling plus a vote over final answers; reported gains included +17.9% on GSM8K and +11.0% on SVAMP.
How to read it: Look at how accuracy grows with the number of sampled paths, and notice the method needs answers that can be compared exactly.
Tree of Thoughts: Deliberate Problem Solving with Large Language Models
Shunyu Yao, Dian Yu et al. · 2023 · NeurIPS 2023
Made classical search explicit around a language model: propose partial solutions, evaluate them, and backtrack instead of committing left to right.
- Problem
- Token-by-token, left-to-right generation cannot look ahead or undo an early choice on tasks that need exploration.
- What was new
- Search (breadth- or depth-first) over “thoughts”, coherent chunks of text, with the model also evaluating states; on Game of 24, GPT-4 went from 4% with chain of thought to 74%.
How to read it: For each task, write down what a state is, who scores it and how many model calls a solution costs; the gains come with a large call budget.
Let's Verify Step by Step
Hunter Lightman, Vineet Kosaraju et al. · 2023
The clearest comparison of rewarding steps versus rewarding final answers, and the source of PRM800K, a widely used set of step-level labels.
- Problem
- A final-answer label cannot say where a derivation went wrong, and it rewards lucky answers reached by invalid steps.
- What was new
- A process-supervised reward model, trained on 800,000 human step labels, selected solutions better than an outcome-supervised one on MATH; the best solved 78% of a representative test subset.
How to read it: Both reward models are used to pick the best of many samples (best-of-N), not to train the generator with RL; keep that setup in mind when generalising.
Scaling Laws for Reward Model Overoptimization
Leo Gao, John Schulman, Jacob Hilton · 2022
Explains why maximizing a learned judge can diverge from the intended outcome.
- Problem
- A policy can exploit errors in a proxy reward model.
- What was new
- Study optimization against a proxy while measuring a synthetic gold reward.
How to read it: Read the synthetic-gold-reward setup before interpreting the scaling curves.
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Zhihong Shao, Peiyi Wang et al. · 2024
Introduced GRPO, the group-relative RL method later used to train DeepSeek-R1 and widely adopted for RL with verifiable rewards.
- Problem
- PPO needs a separate value model as large as the policy, which makes RL on long responses memory-hungry.
- What was new
- A 7B model further pre-trained on 120B math tokens, plus GRPO: a PPO variant that baselines each response against a group of responses to the same prompt instead of a learned critic.
How to read it: For GRPO, go to the RL section and compare its objective with PPO's term by term; the data pipeline sections are a separate, also useful, story.
STaR: Bootstrapping Reasoning With Reasoning
Eric Zelikman, Yuhuai Wu et al. · 2022 · NeurIPS 2022
An early, clear version of the loop behind synthetic reasoning data: a model generates rationales, keeps the ones that reach correct answers, and trains on them.
- Problem
- Training a model to write rationales seemed to need large human-written rationale datasets.
- What was new
- The Self-Taught Reasoner loop: generate rationales, retry failures with the correct answer as a hint (“rationalization”), fine-tune on rationales that ended correctly, repeat.
How to read it: Notice the filter is the final answer only, so a rationale that reaches the right answer by a flawed route is kept; compare Chapter 14's lucky answer.
Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Charlie Snell, Jaehoon Lee et al. · 2024
Treated inference compute as a budget to allocate, and showed the right allocation depends on how hard the prompt is.
- Problem
- Earlier results on spending more inference compute were mixed, with little guidance on which method to use when.
- What was new
- Compared searching against process-based verifiers with revising answers sequentially; allocating compute per prompt by difficulty was over 4× more efficient than best-of-N, and in a FLOPs-matched comparison could beat a 14× larger model on problems the small model sometimes solved.
How to read it: Keep the condition attached to the headline: test-time compute beat a larger model only where the smaller model already had non-trivial success.
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
DeepSeek-AI et al. · 2025
An openly released reasoning model, with a detailed account of training long chains of reasoning mainly through reinforcement learning on verifiable problems.
- Problem
- How reasoning models were trained was largely undisclosed.
- What was new
- Large-scale RL with rule-based rewards on math and code, plus distillation of the resulting reasoning into smaller open models.
How to read it: Read the R1-Zero and R1 sections separately: the first is RL straight from a base model, the second a multi-stage recipe with cold-start data, and the distillation results are a third story.
Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting
Miles Turpin, Julian Michael et al. · 2023 · NeurIPS 2023
Showed with controlled experiments that a chain of thought can omit what actually drove the answer, so a plausible explanation is not evidence of faithfulness.
- Problem
- It is tempting to read a model's step-by-step explanation as the reason for its answer.
- What was new
- Biasing features, such as always making the few-shot answer “(A)”, changed answers without being mentioned; explanations rationalised the biased answers, and accuracy fell by up to 36% on 13 BIG-Bench Hard tasks.
How to read it: The experimental design is the lesson: change an input the explanation never mentions and see whether the answer moves.
Watch
Andrej Karpathy
Deep Dive into LLMs like ChatGPT
A long, general-audience walk through the whole pipeline behind a chat model, from internet text to tokens to pretraining to post-training.
Covers: The full training stack behind a chat model, from internet text and tokens to a pretrained base model and the post-training that turns it into an assistant.
What came next?
Chapter 15
Multimodal AI
The world isn't made of text.
This chapter is being written.