Concept · Chapter 9: How an LLM Is Actually Built
Compute-Optimal Training (Kaplan vs Chinchilla)
For a fixed compute budget, the lowest loss comes from growing parameters and training tokens together, about 20 tokens per parameter by the Chinchilla analysis, but labs deliberately train smaller models on far more tokens when the model will serve many users.
The problem
C ≈ 6ND means a budget can buy a big model trained briefly or a small model trained long. Which split gives the best model, and does the answer change with scale?
The solution
Train many models at several budgets, fit how loss depends on N and D, and find the minimum along each line of constant compute (an IsoFLOP curve); then adjust for the cost of serving the model.
The consequence
Chinchilla (2022) showed that most large models of its time were undertrained; models since have been trained on far more tokens, often well beyond the training-optimal ratio, because inference is the bigger bill.
You should understand first
- Text as Data
- Vectors
- One-Hot Encoding
- Tokenization
- Probability and Distributions
- Conditional Probability and Bayes' Theorem
- Probability of Sequences
- Language Modeling
- Dot Product
- Embeddings
- Attention
- Softmax
- Self-Attention
- Causal Masking
- Entropy
- Loss Functions
- Cross-Entropy Loss
- Autoregressive Next-Token Prediction
- Pretraining at Scale
- Parameters, Tokens and Context Windows
- Scaling Laws (Preview)
- Matrix Multiplication
- Derivatives and Gradients
- The Chain Rule
- The Turing Test
- Symbolic AI
- Logic and Rules
- Expert Systems
- Knowledge Representation
- The Knowledge-Acquisition Bottleneck
- From Rules to Learning
- Supervised, Unsupervised and Self-Supervised Learning
- Features, Labels and Tasks
- Gradient Descent
- Linear Regression
- Logistic Regression
- The Perceptron
- Activation Functions
- The Artificial Neuron
- Multilayer Perceptron (MLP)
- The Forward Pass
- Computational Graphs and Autodiff
- Backpropagation
- The Compute Budget (C ≈ 6ND)
- Compute-Optimal Training (Kaplan vs Chinchilla)
The trade-off
With the budget C fixed, : double the model and you halve its training tokens. Too big, and the model is undertrained. Too small, and extra tokens barely help a model with too little capacity. Somewhere between is a minimum. Plot loss against N along a fixed budget (an IsoFLOP curve) and it is U-shaped.
Kaplan, then Chinchilla
Kaplan and colleagues (2020) estimated that the optimal model size grows as about C^0.73 and the data as C^0.27 Established: put most new compute into parameters. That is roughly what GPT-3 did, with 175B parameters trained on 300B tokens.
Hoffmann and colleagues (2022) trained over 400 models and found, by three different methods, that parameters and data should grow about equally, with exponents near 0.5 each Established. To test it they trained Chinchilla, 70B parameters on 1.4 trillion tokens, with the same compute as the 280B Gopher (trained on 300B tokens), and Chinchilla outperformed Gopher Established. That is 20 tokens per parameter, a ratio that became a rule of thumb.
Tiny example. A budget of 10²² FLOPs at 20 tokens per parameter: 6 × N × 20N = 10²² gives N ≈ 9 billion parameters and D ≈ 180 billion tokens.
The fit is a model too
Chinchilla also fitted a formula . In 2024, Besiroglu and colleagues reconstructed the data from the paper's figure and found that the printed coefficients were inconsistent with the paper's other two methods and had implausibly narrow confidence intervals; their re-fit agreed with about 20 tokens per parameter Established. The lab below lets you switch between the two fits. At Chinchilla's budget the re-fit recommends about 72B parameters, the printed fit about 32B.
Try it
Split a fixed number of FLOPs between model size and training tokens with a published scaling law, place real models on the map, and see how serving costs change the answer.
Optimal for whom?
"Compute-optimal" minimises training cost for a given loss. But a deployed model costs about 2N FLOPs for every token it generates, for as long as it is used. Sardana and colleagues extended the Chinchilla laws with inference demand and found that, with large expected demand (around a billion requests), the cheapest model to train and serve is smaller and trained on more data than Chinchilla-optimal Established. Llama 3's report says its flagship is approximately compute-optimal, while its smaller models were trained for much longer than compute-optimal and perform better than compute-optimal models at the same inference budget Established.
Labs now fit these laws on their own data. Meta's fit, extrapolated to Llama 3's budget of 3.8 × 10²⁵ FLOPs, suggested a 402B model on 16.55T tokens; noting that the IsoFLOP curves flatten near their minimum at large budgets, they chose 405B Established.
Why should I care?
As a researcher
It is the clearest example of an empirical law driving billion-dollar decisions, and of how fragile such fits can be: a 2024 replication found Chinchilla's printed fit inconsistent with its own results.
As an engineer
It explains model menus: why there are 8B models trained on 15 trillion tokens, and why a smaller, longer-trained model can be the cheaper one to deploy.
Modern systems that depend on it
- model sizing
- data requirements
- inference cost planning
Historical context
Before
Kaplan et al. (2020) recommended spending most extra compute on parameters (N_opt ∝ C^0.73), so models grew far faster than their datasets: GPT-3 saw under 2 tokens per parameter.
After
Parameters and tokens scaled about equally (N_opt ∝ C^0.5); Chinchilla-style budgets of about 20 tokens per parameter, and deliberate 'overtraining' of models meant for heavy use.
Used today
Labs fit their own scaling laws on their own data to size models (Llama 3 chose 405B this way) and then train smaller variants far longer than compute-optimal.
What to remember
- IsoFLOP: fix C, vary N (so D = C / 6N), find the N with the lowest loss.
- Kaplan 2020: N_opt ∝ C^0.73. Chinchilla 2022: N_opt ∝ C^0.5, D_opt ∝ C^0.5.
- Rule of thumb: about 20 training tokens per parameter (Chinchilla: 70B on 1.4T).
- Chinchilla's printed parametric fit implies a very different ratio; a 2024 re-fit restores about 20.
- If a model will serve many tokens, a smaller model trained longer is cheaper overall.
Key papers
Scaling Laws for Neural Language Models
Jared Kaplan, Sam McCandlish et al. · 2020
Found that language-model loss falls as a smooth power law in parameters, data and compute — making model scale something you could plan.
Training Compute-Optimal Large Language Models
Jordan Hoffmann, Sebastian Borgeaud et al. · 2022 · NeurIPS 2022
Showed that many large models were undertrained: for a fixed compute budget, parameters and training tokens should grow roughly in proportion.
Scaling Data-Constrained Language Models
Niklas Muennighoff, Alexander M. Rush et al. · 2023
Asked what happens when high-quality text runs out: how much is repeated data worth?
Beyond Chinchilla-Optimal: Accounting for Inference in Language Model Scaling Laws
Nikhil Sardana, Jacob Portes et al. · 2023
Formalised why labs train models far past the 'compute-optimal' point: the model will be run billions of times.
Chinchilla Scaling: A replication attempt
Tamay Besiroglu, Ege Erdil et al. · 2024
Re-fitted Chinchilla's scaling law from its published data and found the printed coefficients inconsistent with the paper's own conclusions: a lesson in checking influential numbers.
How to read it: Short and readable; a good model of how to replicate a result from a figure.
The Llama 3 Herd of Models
Aaron Grattafiori, Abhimanyu Dubey et al. · 2024
The most complete public account of building a frontier-scale model end to end: data pipeline, scaling-law experiments, 16,384-GPU training, failures and all.
How to read it: It is 90+ pages. For this chapter read Section 3 (pre-training) only: data, scaling laws, infrastructure and the training recipe.
Watch
Stanford Online
Stanford CS336 Language Modeling from Scratch | Spring 2025 | Lecture 9: Scaling laws 1
Goes beyond 'bigger is better' to how scaling laws are fitted and used to make decisions.