Concept · Chapter 9: How an LLM Is Actually Built
The Compute Budget (C ≈ 6ND)
Training a dense Transformer with N parameters on D tokens costs about 6ND floating-point operations, which, divided by what a GPU actually sustains, turns a model plan into GPU-days and money.
The problem
Before spending months and millions on a run, a team needs to know how much computation it requires and how long it will take on the hardware it has.
The solution
Count FLOPs per token (about 2N for the forward pass and 4N for the backward pass), multiply by tokens, and divide by the throughput each GPU really achieves (its model FLOPs utilisation times its peak).
The consequence
Compute became the common currency of AI progress: budgets, scaling laws, policy thresholds and comparisons between models are all stated in training FLOPs.
You should understand first
- Text as Data
- Vectors
- One-Hot Encoding
- Tokenization
- Probability and Distributions
- Conditional Probability and Bayes' Theorem
- Probability of Sequences
- Language Modeling
- Dot Product
- Embeddings
- Attention
- Softmax
- Self-Attention
- Causal Masking
- Entropy
- Loss Functions
- Cross-Entropy Loss
- Autoregressive Next-Token Prediction
- Pretraining at Scale
- Parameters, Tokens and Context Windows
- Matrix Multiplication
- Derivatives and Gradients
- The Chain Rule
- The Turing Test
- Symbolic AI
- Logic and Rules
- Expert Systems
- Knowledge Representation
- The Knowledge-Acquisition Bottleneck
- From Rules to Learning
- Supervised, Unsupervised and Self-Supervised Learning
- Features, Labels and Tasks
- Gradient Descent
- Linear Regression
- Logistic Regression
- The Perceptron
- Activation Functions
- The Artificial Neuron
- Multilayer Perceptron (MLP)
- The Forward Pass
- Computational Graphs and Autodiff
- Backpropagation
- The Compute Budget (C ≈ 6ND)
Where 6 comes from
Most of a Transformer's work is multiplying activations by weight matrices. Each weight is used once per token in the forward pass, as one multiply and one add: about 2N FLOPs per token for N parameters. The backward pass computes gradients with respect to both the activations and the weights, about twice as much again: 4N. Kaplan and colleagues therefore estimated training compute as C ≈ 6N floating-point operations per training token Established:
The estimate leaves out attention's own cost, which grows with the context length; it is small whenever the model's width is large relative to the context Established.
Checking it against real runs
| Model | N | D | 6ND | Reported |
|---|---|---|---|---|
| GPT-3 | 175B | 300B | 3.15 × 10²³ | 3.14 × 10²³ Established |
| Chinchilla | 70B | 1.4T | 5.9 × 10²³ | 5.76 × 10²³ (the Gopher budget) Established |
| Llama 3 405B | 405B | 15.6T | 3.79 × 10²⁵ | 3.8 × 10²⁵ Established |
From FLOPs to days
GPUs never run at their peak. Model FLOPs utilisation (MFU) is the useful model FLOPs per second divided by the hardware's peak. PaLM 540B reached 46.2% MFU Established; Llama 3 405B reached 38–43%, about 400 TFLOP/s per H100 on 16,384 GPUs Established.
Tiny example. 3.8 × 10²⁵ FLOPs ÷ (16,384 GPUs × 4 × 10¹⁴ FLOP/s) ≈ 5.8 × 10⁶ seconds ≈ 67 days of uninterrupted training at that rate. This is our arithmetic, not a reported duration; real runs also lose time to failures and restarts.
Why should I care?
As a researcher
Every scaling-law plot has compute on its x-axis; a claim about efficiency is a claim about loss per FLOP.
As an engineer
6ND plus a utilisation figure is the napkin maths for whether a training or fine-tuning job is hours, days or impossible on your hardware.
Modern systems that depend on it
- scaling laws
- compute-optimal training
- cluster sizing
- cost estimates
Historical context
Before
Training cost was described by GPU type and wall-clock time, which made runs hard to compare.
After
Runs are described by total FLOPs and achieved utilisation; 6ND is the standard first estimate.
Used today
Model reports quote training FLOPs (Llama 3 405B: 3.8 × 10²⁵); MFU is the standard efficiency metric for large runs.
What to remember
- C ≈ 6ND: about 2N FLOPs per token forward, 4N backward.
- GPT-3: 6 × 175B × 300B ≈ 3.1 × 10²³ FLOPs. Llama 3 405B: 6 × 405B × 15.6T ≈ 3.8 × 10²⁵.
- MFU = achieved model FLOPs per second ÷ hardware peak. Large runs reach roughly 40–50%.
- Days ≈ C ÷ (GPUs × FLOP/s per GPU × 86,400).
- 6ND ignores attention's sequence-length term, which matters only for long contexts relative to model width.
Key papers
Language Models are Few-Shot Learners
Tom B. Brown, Benjamin Mann et al. · 2020 · NeurIPS 2020
GPT-3 (175B parameters) showed that a large enough language model can perform new tasks from a few examples in its prompt, without any gradient updates.
How to read it: 75 pages. Sections 1–2 and Figure 1.2 carry the core idea; Section 6 on broader impacts is worth reading too.
Scaling Laws for Neural Language Models
Jared Kaplan, Sam McCandlish et al. · 2020
Found that language-model loss falls as a smooth power law in parameters, data and compute — making model scale something you could plan.
PaLM: Scaling Language Modeling with Pathways
Aakanksha Chowdhery, Sharan Narang et al. · 2022
A 540-billion-parameter model whose report defined model FLOPs utilization (MFU) and described, unusually frankly, the loss spikes of a very large run.
How to read it: Sections 4 (training infrastructure) and 5.1 (training instability) are the parts for this chapter.
The Llama 3 Herd of Models
Aaron Grattafiori, Abhimanyu Dubey et al. · 2024
The most complete public account of building a frontier-scale model end to end: data pipeline, scaling-law experiments, 16,384-GPU training, failures and all.
How to read it: It is 90+ pages. For this chapter read Section 3 (pre-training) only: data, scaling laws, infrastructure and the training recipe.
Watch
Stanford Online
Stanford CS336 Language Modeling from Scratch | Spring 2025 | Lec. 2: Pytorch, Resource Accounting
Teaches the habit this chapter is built on: napkin maths for memory (bytes per parameter) and compute (FLOPs) before you train anything.
Stanford Online
Stanford CS336 Language Modeling from Scratch | Spring 2025 | Lecture 9: Scaling laws 1
Goes beyond 'bigger is better' to how scaling laws are fitted and used to make decisions.