Skip to content
Road to Intelligence

Concept · Chapter 14: Reasoning Models

Inference-Time Scaling

Must knowKnow well12 minDifficulty

Inference-time scaling is improving answers by spending more computation per request, on longer reasoning, more samples, search or verification, and it pays off only when that computation is aimed by a useful selector and matched to the problem's difficulty.

The problem

A fixed, short answer budget leaves hard problems unsolved, while a large fixed budget wastes compute and latency on easy ones.

The solution

Treat per-request compute as a budget split across depth (longer chains), breadth (more samples) and verification, and allocate it according to how hard the request is.

The consequence

Accuracy can rise substantially with inference compute on suitable tasks, and the cost of serving a model now depends on how much each request is allowed to think.

Three budgets

Every inference-time method spends compute on one or more of:

  • Depth: a longer single attempt, more steps of chain of thought, more moves in a search.
  • Breadth: more independent attempts, as in self-consistency and best-of-N.
  • Verification: checking, scoring or revising candidates, as with verifiers.

They fail differently. Breadth cannot recover a solution the generator never produces. Verification cannot select a correct answer absent from the batch. Depth cannot help if the chain heads the wrong way and nothing checks it. The search lab shows the first two directly: width 3 rescues the route to 19 at six moves, and no width finds it in five.

The evidence

OpenAI's o1 announcement reported that the model's performance improved consistently with more reinforcement-learning training compute and with more time spent thinking at test time. Established It gave the trend as plots without the underlying training details, so treat it as a reported observation for one model family.

Snell and colleagues compared searching against process-based verifiers with letting the model revise its answers, and found the most effective method depended on problem difficulty; choosing per prompt was more than 4× more efficient than a best-of-N baseline, and in a FLOPs-matched comparison test-time compute could outperform a 14× larger model on problems where the smaller model already had non-trivial success. Established

The condition in that last sentence is the point: extra inference compute helps most where the model is close to solving the problem already.

What a request actually costs

For a reasoning request, count:

  • prompt tokens (prefill, once per sample unless cached);
  • generated tokens, including hidden reasoning tokens, which are usually billed;
  • number of samples;
  • verifier or tool calls;
  • latency, which grows with the longest sequential chain.

Tiny example. One attempt of 2,000 reasoning tokens versus 8 parallel attempts of 500 tokens each. The parallel option generates twice the tokens (4,000) but each sequence is a quarter as long, so wall-clock latency can be lower if the hardware runs the samples together. The KV cache is the other cost: eight concurrent sequences hold eight caches.

Allocate by difficulty

An easy request answered in 20 tokens and a hard one given 20,000 should not share a budget. A budget policy decides:

  1. how to estimate difficulty (a classifier, the model's own uncertainty, disagreement among a few cheap samples);
  2. which axis to spend on first;
  3. when to stop, and what to report when the budget runs out without a verified answer.
Evaluate a system's budget policy, not just its model: which requests get more work, when it stops, and how it behaves at the limit. Interpretation

Reading results

When a paper or product reports a gain from more inference compute, ask: compute measured how (tokens, FLOPs, dollars)? Selection by what (vote, learned scorer, oracle)? Compared at matched compute against what baseline (more samples from a bigger model, a cheaper model with more samples)? And on which difficulty range?

Mini experiment

In the consensus lab, at 20% fresh accuracy, compare the verifier's success with 4, 16 and 32 candidates over a few seeds. Then estimate the cost multiplier of each budget relative to 4. Where would you stop if each extra candidate cost one cent and a wrong answer cost ten?

Why should I care?

As a researcher

Test-time compute is a second scaling axis alongside training compute, and how to allocate it per prompt is an active research question.

As an engineer

Reasoning modes change your latency, cost per request and KV-cache memory, so you need a budget policy, not just a model choice.

Modern systems that depend on it

  • reasoning-effort settings in model APIs
  • adaptive routing between fast and slow models
  • evaluation at matched compute

Historical context

Before

A model's quality was fixed at training time; inference cost was roughly proportional to the length of a short answer.

After

Quality depends on how much inference work a request is given, and serving a model means choosing a budget policy.

Used today

Model APIs expose reasoning-effort or thinking-budget settings, and reported benchmark results often state the sample count or effort level.

What to remember

  • Three budgets: depth (longer chains), breadth (more samples), verification (checking and selecting).
  • Breadth cannot help if the generator never succeeds; a better verifier cannot select an answer that is not there.
  • Snell et al.: the best strategy depends on difficulty; per-prompt allocation beat best-of-N by more than 4× in efficiency.
  • o1 (reported by OpenAI): performance improved with both more RL training compute and more thinking time.
  • Parallel samples cut latency relative to one long chain but not total compute.
  • Compare systems at matched compute and latency, with the selection rule stated.

Key papers

Important

Training Verifiers to Solve Math Word Problems

Karl Cobbe, Vineet Kosaraju et al. · 2021

Introduced GSM8K, the grade-school math benchmark used for years afterwards, and the generate-many-then-verify recipe that best-of-N selection still follows.

How to read it: Read the sections on the verifier and on how test performance changes with the number of sampled completions; that trade-off is Chapter 14's selection problem.

~35 min readarXiv:2110.14168✓ verified 2026-10-06
Essential

Self-Consistency Improves Chain of Thought Reasoning in Language Models

Xuezhi Wang, Jason Wei et al. · 2022 · ICLR 2023

Turned extra inference compute into accuracy without any training: sample many reasoning paths and keep the answer they most often reach.

How to read it: Look at how accuracy grows with the number of sampled paths, and notice the method needs answers that can be compared exactly.

~35 min readarXiv:2203.11171✓ verified 2026-10-06
Essential

Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Charlie Snell, Jaehoon Lee et al. · 2024

Treated inference compute as a budget to allocate, and showed the right allocation depends on how hard the prompt is.

How to read it: Keep the condition attached to the headline: test-time compute beat a larger model only where the smaller model already had non-trivial success.

~1 h readarXiv:2408.03314✓ verified 2026-10-06