Concept · Chapter 14: Reasoning Models
Training vs Inference Compute
Training compute changes a model's parameters once, for every future request; inference compute is spent on one request with the parameters held fixed, and nothing it learns survives unless someone saves it.
The problem
“Spend more compute” is ambiguous: a bigger training run, a longer answer and fifty sampled attempts all cost compute, but they change different things and are paid for at different times.
The solution
Keep two ledgers. Training compute buys a new checkpoint; inference compute buys work on the current request: intermediate tokens, extra samples, search and verification.
The consequence
Reasoning systems can trade between the ledgers, for example by training a model to use inference compute well, or by distilling expensive inference into a cheaper student, and the trade only pays off at matched quality.
You should understand first
- Text as Data
- Probability and Distributions
- Conditional Probability and Bayes' Theorem
- Probability of Sequences
- Language Modeling
- Vectors
- Dot Product
- Embeddings
- Attention
- Softmax
- Self-Attention
- Causal Masking
- Entropy
- Loss Functions
- Cross-Entropy Loss
- Autoregressive Next-Token Prediction
- Pretraining at Scale
- Training vs Inference Compute
Two ledgers
Think of a data platform. Building a better query planner is a one-off engineering investment that benefits every future query. Giving one expensive query more executors, retries and validation helps that query only. Language models have the same two ledgers.
Training produces parameter updates. Standard inference generates outputs with the parameters fixed, however many forward passes it runs and however many tools it calls. Established A chain of thought, fifty samples or a search tree are all inference: when the request ends, the weights are exactly what they were.
| Spend compute on… | What changes | Who benefits |
|---|---|---|
| Pretraining or fine-tuning | Parameters | Every later request |
| RL with rewards | Parameters (the response distribution) | Every later request |
| Intermediate tokens | This request's context | This request |
| Extra samples, search, verification | This request's candidate set | This request |
| Saving outputs to a database | Application memory, not weights | Later requests that retrieve them |
Inference was never one pass
A common shorthand says a model answers “in one forward pass”. That is true per token, not per answer. A 200-token answer already runs the network about 200 times (after a prefill pass over the prompt, see prefill and decode). Reasoning does not introduce iteration; it changes what the extra iterations produce. A hundred tokens spent writing before the answer give later tokens something to condition on. A hundred tokens of “let me think carefully” may give them nothing.
A tiny ledger
Suppose a training improvement costs and the model then serves requests at an average inference cost :
Try it. A reasoning mode adds 10 units per request. Over a million requests that is 10 million units. If a 2-million-unit training run let the model reach the same quality with 2 extra units per request instead of 10, it would save million units. If quality drops, the comparison is meaningless: the identity is accounting, not a performance law. It also needs one unit throughout (FLOPs, GPU-hours, energy or money) and an honest average over the requests you actually serve.
Where the line blurs
Several techniques in this chapter move work from one ledger to the other:
- RL with verifiable rewards spends training compute so that the model makes better use of inference compute later.
- Distillation spends inference compute on a teacher, once, to produce training data for a cheaper student.
- Retrieval and memory store inference outputs outside the model, so they persist without any weight change.
Snell and colleagues compared extra inference compute with a larger model in a FLOPs-matched evaluation, and found inference compute could win on prompts where the smaller model already had non-trivial success. Established That is a conditional result about one setup, not an exchange rate between the ledgers.
Where it shows up
Model cards and product pages increasingly report results “at high reasoning effort” or with a sample count. When you read one, ask which ledger produced the gain, and what it would cost per request at your traffic. For your own systems, log inference compute per request (prompt tokens, generated tokens, samples, verifier calls) the way you would log query cost in a warehouse.
What to remember
- Training changes weights; inference with fixed weights changes only the context, the candidate set and what the application stores.
- An ordinary answer already runs the model about once per generated token; reasoning changes what those extra passes are used for.
- Total cost ≈ training cost + requests × average inference cost, in one consistent unit.
- A prompt example helps today's answer; only fine-tuning on it changes tomorrow's checkpoint.
- Compare compute allocations at matched quality and latency, never on cost alone.
Key papers
Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Charlie Snell, Jaehoon Lee et al. · 2024
Treated inference compute as a budget to allocate, and showed the right allocation depends on how hard the prompt is.
How to read it: Keep the condition attached to the headline: test-time compute beat a larger model only where the smaller model already had non-trivial success.