Concept · Chapter 14: Reasoning Models
Inference-Time Scaling
Inference-time scaling is improving answers by spending more computation per request, on longer reasoning, more samples, search or verification, and it pays off only when that computation is aimed by a useful selector and matched to the problem's difficulty.
The problem
A fixed, short answer budget leaves hard problems unsolved, while a large fixed budget wastes compute and latency on easy ones.
The solution
Treat per-request compute as a budget split across depth (longer chains), breadth (more samples) and verification, and allocate it according to how hard the request is.
The consequence
Accuracy can rise substantially with inference compute on suitable tasks, and the cost of serving a model now depends on how much each request is allowed to think.
You should understand first
- Text as Data
- Probability and Distributions
- Conditional Probability and Bayes' Theorem
- Probability of Sequences
- Language Modeling
- Vectors
- Dot Product
- Embeddings
- Attention
- Softmax
- Self-Attention
- Causal Masking
- Entropy
- Loss Functions
- Cross-Entropy Loss
- Autoregressive Next-Token Prediction
- Pretraining at Scale
- Training vs Inference Compute
- The Turing Test
- Symbolic AI
- Search
- GPT-1 → GPT-2 → GPT-3
- In-Context Learning
- Chain of Thought
- Search over Reasoning Steps
- Logic and Rules
- Expert Systems
- Knowledge Representation
- The Knowledge-Acquisition Bottleneck
- From Rules to Learning
- Supervised, Unsupervised and Self-Supervised Learning
- Decoding: Greedy, Temperature, Top-k, Top-p
- Learning from Comparisons
- Features, Labels and Tasks
- Derivatives and Gradients
- Gradient Descent
- Linear Regression
- Logistic Regression
- Reward Models
- Expected Value and Variance
- Sampling and Uncertainty
- Self-Consistency
- Verifiers and Best-of-N
- Inference-Time Scaling
Three budgets
Every inference-time method spends compute on one or more of:
- Depth: a longer single attempt, more steps of chain of thought, more moves in a search.
- Breadth: more independent attempts, as in self-consistency and best-of-N.
- Verification: checking, scoring or revising candidates, as with verifiers.
They fail differently. Breadth cannot recover a solution the generator never produces. Verification cannot select a correct answer absent from the batch. Depth cannot help if the chain heads the wrong way and nothing checks it. The search lab shows the first two directly: width 3 rescues the route to 19 at six moves, and no width finds it in five.
The evidence
OpenAI's o1 announcement reported that the model's performance improved consistently with more reinforcement-learning training compute and with more time spent thinking at test time. Established It gave the trend as plots without the underlying training details, so treat it as a reported observation for one model family.
Snell and colleagues compared searching against process-based verifiers with letting the model revise its answers, and found the most effective method depended on problem difficulty; choosing per prompt was more than 4× more efficient than a best-of-N baseline, and in a FLOPs-matched comparison test-time compute could outperform a 14× larger model on problems where the smaller model already had non-trivial success. EstablishedThe condition in that last sentence is the point: extra inference compute helps most where the model is close to solving the problem already.
What a request actually costs
For a reasoning request, count:
- prompt tokens (prefill, once per sample unless cached);
- generated tokens, including hidden reasoning tokens, which are usually billed;
- number of samples;
- verifier or tool calls;
- latency, which grows with the longest sequential chain.
Tiny example. One attempt of 2,000 reasoning tokens versus 8 parallel attempts of 500 tokens each. The parallel option generates twice the tokens (4,000) but each sequence is a quarter as long, so wall-clock latency can be lower if the hardware runs the samples together. The KV cache is the other cost: eight concurrent sequences hold eight caches.
Allocate by difficulty
An easy request answered in 20 tokens and a hard one given 20,000 should not share a budget. A budget policy decides:
- how to estimate difficulty (a classifier, the model's own uncertainty, disagreement among a few cheap samples);
- which axis to spend on first;
- when to stop, and what to report when the budget runs out without a verified answer.
Reading results
When a paper or product reports a gain from more inference compute, ask: compute measured how (tokens, FLOPs, dollars)? Selection by what (vote, learned scorer, oracle)? Compared at matched compute against what baseline (more samples from a bigger model, a cheaper model with more samples)? And on which difficulty range?
Mini experiment
In the consensus lab, at 20% fresh accuracy, compare the verifier's success with 4, 16 and 32 candidates over a few seeds. Then estimate the cost multiplier of each budget relative to 4. Where would you stop if each extra candidate cost one cent and a wrong answer cost ten?
Why should I care?
As a researcher
Test-time compute is a second scaling axis alongside training compute, and how to allocate it per prompt is an active research question.
As an engineer
Reasoning modes change your latency, cost per request and KV-cache memory, so you need a budget policy, not just a model choice.
Modern systems that depend on it
- reasoning-effort settings in model APIs
- adaptive routing between fast and slow models
- evaluation at matched compute
Historical context
Before
A model's quality was fixed at training time; inference cost was roughly proportional to the length of a short answer.
After
Quality depends on how much inference work a request is given, and serving a model means choosing a budget policy.
Used today
Model APIs expose reasoning-effort or thinking-budget settings, and reported benchmark results often state the sample count or effort level.
What to remember
- Three budgets: depth (longer chains), breadth (more samples), verification (checking and selecting).
- Breadth cannot help if the generator never succeeds; a better verifier cannot select an answer that is not there.
- Snell et al.: the best strategy depends on difficulty; per-prompt allocation beat best-of-N by more than 4× in efficiency.
- o1 (reported by OpenAI): performance improved with both more RL training compute and more thinking time.
- Parallel samples cut latency relative to one long chain but not total compute.
- Compare systems at matched compute and latency, with the selection rule stated.
Key papers
Training Verifiers to Solve Math Word Problems
Karl Cobbe, Vineet Kosaraju et al. · 2021
Introduced GSM8K, the grade-school math benchmark used for years afterwards, and the generate-many-then-verify recipe that best-of-N selection still follows.
How to read it: Read the sections on the verifier and on how test performance changes with the number of sampled completions; that trade-off is Chapter 14's selection problem.
Self-Consistency Improves Chain of Thought Reasoning in Language Models
Xuezhi Wang, Jason Wei et al. · 2022 · ICLR 2023
Turned extra inference compute into accuracy without any training: sample many reasoning paths and keep the answer they most often reach.
How to read it: Look at how accuracy grows with the number of sampled paths, and notice the method needs answers that can be compared exactly.
Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Charlie Snell, Jaehoon Lee et al. · 2024
Treated inference compute as a budget to allocate, and showed the right allocation depends on how hard the prompt is.
How to read it: Keep the condition attached to the headline: test-time compute beat a larger model only where the smaller model already had non-trivial success.