Skip to content
Road to Intelligence

Concept · Chapter 14: Reasoning Models

Training vs Inference Compute

Must knowKnow well9 minDifficulty

Training compute changes a model's parameters once, for every future request; inference compute is spent on one request with the parameters held fixed, and nothing it learns survives unless someone saves it.

The problem

“Spend more compute” is ambiguous: a bigger training run, a longer answer and fifty sampled attempts all cost compute, but they change different things and are paid for at different times.

The solution

Keep two ledgers. Training compute buys a new checkpoint; inference compute buys work on the current request: intermediate tokens, extra samples, search and verification.

The consequence

Reasoning systems can trade between the ledgers, for example by training a model to use inference compute well, or by distilling expensive inference into a cheaper student, and the trade only pays off at matched quality.

Two ledgers

Think of a data platform. Building a better query planner is a one-off engineering investment that benefits every future query. Giving one expensive query more executors, retries and validation helps that query only. Language models have the same two ledgers.

Training produces parameter updates. Standard inference generates outputs with the parameters fixed, however many forward passes it runs and however many tools it calls. Established A chain of thought, fifty samples or a search tree are all inference: when the request ends, the weights are exactly what they were.

Spend compute on…What changesWho benefits
Pretraining or fine-tuningParametersEvery later request
RL with rewardsParameters (the response distribution)Every later request
Intermediate tokensThis request's contextThis request
Extra samples, search, verificationThis request's candidate setThis request
Saving outputs to a databaseApplication memory, not weightsLater requests that retrieve them

Inference was never one pass

A common shorthand says a model answers “in one forward pass”. That is true per token, not per answer. A 200-token answer already runs the network about 200 times (after a prefill pass over the prompt, see prefill and decode). Reasoning does not introduce iteration; it changes what the extra iterations produce. A hundred tokens spent writing 6×7=426 \times 7 = 42 before the answer give later tokens something to condition on. A hundred tokens of “let me think carefully” may give them nothing.

A tiny ledger

Suppose a training improvement costs CtrainC_{\text{train}} and the model then serves QQ requests at an average inference cost CinferC_{\text{infer}}:

Ctotal=Ctrain+Q Cinfer.C_{\text{total}} = C_{\text{train}} + Q \, C_{\text{infer}}.

Try it. A reasoning mode adds 10 units per request. Over a million requests that is 10 million units. If a 2-million-unit training run let the model reach the same quality with 2 extra units per request instead of 10, it would save 8×106−2×106=68 \times 10^6 - 2 \times 10^6 = 6 million units. If quality drops, the comparison is meaningless: the identity is accounting, not a performance law. It also needs one unit throughout (FLOPs, GPU-hours, energy or money) and an honest average over the requests you actually serve.

Where the line blurs

Several techniques in this chapter move work from one ledger to the other:

  • RL with verifiable rewards spends training compute so that the model makes better use of inference compute later.
  • Distillation spends inference compute on a teacher, once, to produce training data for a cheaper student.
  • Retrieval and memory store inference outputs outside the model, so they persist without any weight change.

Snell and colleagues compared extra inference compute with a larger model in a FLOPs-matched evaluation, and found inference compute could win on prompts where the smaller model already had non-trivial success. Established That is a conditional result about one setup, not an exchange rate between the ledgers.

Where it shows up

Model cards and product pages increasingly report results “at high reasoning effort” or with a sample count. When you read one, ask which ledger produced the gain, and what it would cost per request at your traffic. For your own systems, log inference compute per request (prompt tokens, generated tokens, samples, verifier calls) the way you would log query cost in a warehouse.

What to remember

  • Training changes weights; inference with fixed weights changes only the context, the candidate set and what the application stores.
  • An ordinary answer already runs the model about once per generated token; reasoning changes what those extra passes are used for.
  • Total cost ≈ training cost + requests × average inference cost, in one consistent unit.
  • A prompt example helps today's answer; only fine-tuning on it changes tomorrow's checkpoint.
  • Compare compute allocations at matched quality and latency, never on cost alone.

Key papers

Essential

Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Charlie Snell, Jaehoon Lee et al. · 2024

Treated inference compute as a budget to allocate, and showed the right allocation depends on how hard the prompt is.

How to read it: Keep the condition attached to the headline: test-time compute beat a larger model only where the smaller model already had non-trivial success.

~1 h readarXiv:2408.03314✓ verified 2026-10-06