Skip to content
Road to Intelligence

Concept · Chapter 9: How an LLM Is Actually Built

Compute-Optimal Training (Kaplan vs Chinchilla)

Must knowKnow well16 minDifficulty

For a fixed compute budget, the lowest loss comes from growing parameters and training tokens together, about 20 tokens per parameter by the Chinchilla analysis, but labs deliberately train smaller models on far more tokens when the model will serve many users.

The problem

C ≈ 6ND means a budget can buy a big model trained briefly or a small model trained long. Which split gives the best model, and does the answer change with scale?

The solution

Train many models at several budgets, fit how loss depends on N and D, and find the minimum along each line of constant compute (an IsoFLOP curve); then adjust for the cost of serving the model.

The consequence

Chinchilla (2022) showed that most large models of its time were undertrained; models since have been trained on far more tokens, often well beyond the training-optimal ratio, because inference is the bigger bill.

The trade-off

With the budget C fixed, D=C/6ND = C/6N: double the model and you halve its training tokens. Too big, and the model is undertrained. Too small, and extra tokens barely help a model with too little capacity. Somewhere between is a minimum. Plot loss against N along a fixed budget (an IsoFLOP curve) and it is U-shaped.

Kaplan, then Chinchilla

Kaplan and colleagues (2020) estimated that the optimal model size grows as about C^0.73 and the data as C^0.27 Established: put most new compute into parameters. That is roughly what GPT-3 did, with 175B parameters trained on 300B tokens.

Hoffmann and colleagues (2022) trained over 400 models and found, by three different methods, that parameters and data should grow about equally, with exponents near 0.5 each Established. To test it they trained Chinchilla, 70B parameters on 1.4 trillion tokens, with the same compute as the 280B Gopher (trained on 300B tokens), and Chinchilla outperformed Gopher Established. That is 20 tokens per parameter, a ratio that became a rule of thumb.

Tiny example. A budget of 10²² FLOPs at 20 tokens per parameter: 6 × N × 20N = 10²² gives N ≈ 9 billion parameters and D ≈ 180 billion tokens.

The fit is a model too

Chinchilla also fitted a formula L(N,D)=E+A/Nα+B/DβL(N,D)=E+A/N^{\alpha}+B/D^{\beta}. In 2024, Besiroglu and colleagues reconstructed the data from the paper's figure and found that the printed coefficients were inconsistent with the paper's other two methods and had implausibly narrow confidence intervals; their re-fit agreed with about 20 tokens per parameter Established. The lab below lets you switch between the two fits. At Chinchilla's budget the re-fit recommends about 72B parameters, the printed fit about 32B.

Try it

Spend a Compute Budget

Split a fixed number of FLOPs between model size and training tokens with a published scaling law, place real models on the map, and see how serving costs change the answer.

Know well9 min

Optimal for whom?

"Compute-optimal" minimises training cost for a given loss. But a deployed model costs about 2N FLOPs for every token it generates, for as long as it is used. Sardana and colleagues extended the Chinchilla laws with inference demand and found that, with large expected demand (around a billion requests), the cheapest model to train and serve is smaller and trained on more data than Chinchilla-optimal Established. Llama 3's report says its flagship is approximately compute-optimal, while its smaller models were trained for much longer than compute-optimal and perform better than compute-optimal models at the same inference budget Established.

Labs now fit these laws on their own data. Meta's fit, extrapolated to Llama 3's budget of 3.8 × 10²⁵ FLOPs, suggested a 402B model on 16.55T tokens; noting that the IsoFLOP curves flatten near their minimum at large budgets, they chose 405B Established.

Why should I care?

As a researcher

It is the clearest example of an empirical law driving billion-dollar decisions, and of how fragile such fits can be: a 2024 replication found Chinchilla's printed fit inconsistent with its own results.

As an engineer

It explains model menus: why there are 8B models trained on 15 trillion tokens, and why a smaller, longer-trained model can be the cheaper one to deploy.

Modern systems that depend on it

  • model sizing
  • data requirements
  • inference cost planning

Historical context

Before

Kaplan et al. (2020) recommended spending most extra compute on parameters (N_opt ∝ C^0.73), so models grew far faster than their datasets: GPT-3 saw under 2 tokens per parameter.

After

Parameters and tokens scaled about equally (N_opt ∝ C^0.5); Chinchilla-style budgets of about 20 tokens per parameter, and deliberate 'overtraining' of models meant for heavy use.

Used today

Labs fit their own scaling laws on their own data to size models (Llama 3 chose 405B this way) and then train smaller variants far longer than compute-optimal.

What to remember

  • IsoFLOP: fix C, vary N (so D = C / 6N), find the N with the lowest loss.
  • Kaplan 2020: N_opt ∝ C^0.73. Chinchilla 2022: N_opt ∝ C^0.5, D_opt ∝ C^0.5.
  • Rule of thumb: about 20 training tokens per parameter (Chinchilla: 70B on 1.4T).
  • Chinchilla's printed parametric fit implies a very different ratio; a 2024 re-fit restores about 20.
  • If a model will serve many tokens, a smaller model trained longer is cheaper overall.

Key papers

Essential

Scaling Laws for Neural Language Models

Jared Kaplan, Sam McCandlish et al. · 2020

Found that language-model loss falls as a smooth power law in parameters, data and compute — making model scale something you could plan.

~1 h readarXiv:2001.08361✓ verified 2026-09-26
Essential

Training Compute-Optimal Large Language Models

Jordan Hoffmann, Sebastian Borgeaud et al. · 2022 · NeurIPS 2022

Showed that many large models were undertrained: for a fixed compute budget, parameters and training tokens should grow roughly in proportion.

~1 h readarXiv:2203.15556✓ verified 2026-09-26
Important

Scaling Data-Constrained Language Models

Niklas Muennighoff, Alexander M. Rush et al. · 2023

Asked what happens when high-quality text runs out: how much is repeated data worth?

~40 min readarXiv:2305.16264✓ verified 2026-10-04
Important

Chinchilla Scaling: A replication attempt

Tamay Besiroglu, Ege Erdil et al. · 2024

Re-fitted Chinchilla's scaling law from its published data and found the printed coefficients inconsistent with the paper's own conclusions: a lesson in checking influential numbers.

How to read it: Short and readable; a good model of how to replicate a result from a figure.

~20 min readarXiv:2404.10102✓ verified 2026-10-04
Essential

The Llama 3 Herd of Models

Aaron Grattafiori, Abhimanyu Dubey et al. · 2024

The most complete public account of building a frontier-scale model end to end: data pipeline, scaling-law experiments, 16,384-GPU training, failures and all.

How to read it: It is 90+ pages. For this chapter read Section 3 (pre-training) only: data, scaling laws, infrastructure and the training recipe.

~2 h readarXiv:2407.21783✓ verified 2026-10-04

Watch