Part III · Large Language Models
Chapter 9
How an LLM Is Actually Built
Data, tokens, GPUs and the laws of scale.
In one sentenceBuilding an LLM is a data, systems and budgeting problem: curate trillions of tokens, split training across thousands of accelerators, and spend compute where scaling laws say it pays.
The problem
The shopping list
Chapter 8 summed up how large language models are made in one sentence: predict the next token, at scale. This chapter is about the words "at scale". What would it actually take to train one?
Most frontier labs no longer say. One exception is Meta's report on Llama 3, which describes training a 405-billion-parameter model in unusual detail, failures included. It makes a good running example. Here is its bill of materials:
| What | Llama 3 405B | Section |
|---|---|---|
| Training text | 15.6 trillion tokens, cleaned from far more | Trillions of tokens; The mix |
| Tokenizer | 128,000-token vocabulary | The tokenizer |
| Training steps | about 1.2 million, each over millions of tokens | The training loop |
| Memory for training state | about 6.5 TB (16 bytes per parameter) | Memory |
| Hardware | up to 16,384 H100 GPUs, 80 GB each | Many GPUs |
| Interruptions | 466 in one 54-day period | When things break |
| Compute | 3.8 × 10²⁵ floating-point operations | Counting the compute |
The token count, vocabulary, GPU count, interruptions and compute are all from the Llama 3 report Established. The 6.5 TB is our arithmetic: 405 billion parameters at 16 bytes each, a figure explained below.
Read down the list and it is less a machine-learning problem than a data-engineering and distributed-systems one. If you build data pipelines for a living, you already know half of this chapter: ETL, deduplication, sharding, checkpoints and capacity planning. The other half is what those words mean when the job is a neural network.
The raw material
Trillions of tokens
The text mostly comes from Common Crawl, a public archive of the web released as snapshots. Raw, it is mostly menus, cookie banners, product listings, spam and copies. A pretraining pipeline extracts the main text from each page, identifies its language, applies cheap rules (too short, too repetitive, too many symbols), scores quality with a classifier, removes unsafe and personal content, and deduplicates.
Most of what goes in does not come out. In the open FineWeb dataset, basic filtering of 96 Common Crawl snapshots left about 36 trillion tokens; deduplication left about 20 trillion; further filters left 15 trillion. An educational-quality classifier, trained on Llama 3's ratings of 460,000 pages, then selected 1.3 trillion of those for FineWeb-Edu Established.
Copies deserve their own step. One 61-word sentence appeared over 60,000 times in the C4 dataset, and models trained after deduplication emitted memorised text ten times less often Established. Exact hashing only catches identical pages, so pipelines compare documents as sets of overlapping word sequences (shingles) and estimate their overlap with compact MinHash signatures. Try the method on eleven made-up web pages. At the default settings, which are FineWeb's, the mirror and the syndicated copy of a news story are caught, as are two shop pages built from one template. A re-posted recipe with a reworded first sentence slips through. Lower the rows per band to catch it.
Try it · toy model
Run the deduplication method web-scale datasets use (shingles, MinHash and banded hashing) on a dozen made-up web pages, and tune what counts as a copy.
Even this has surprises. FineWeb's team found that deduplicating all 96 snapshots against each other made the data worse than deduplicating each snapshot on its own: in the oldest crawls, what survived global deduplication was lower quality than what it removed Established.
The recipe
What goes in the mix
A training set is not one source but a weighted blend. Sampling in proportion to size would bury the small, dense sources (books, encyclopedias, code, papers) under web text. LLaMA (2023) drew 67% of its tokens from Common Crawl and 4.5% from Wikipedia, but went over Wikipedia about 2.45 times and over Common Crawl about 1.1 times Established. Llama 3 chose its mix by training small models on candidate blends and extrapolating with scaling laws, ending at roughly 50% general knowledge, 25% maths and reasoning, 17% code and 8% multilingual text Established.
How often can the good data be repeated? Up to about 4 epochs, repeated data was almost as valuable as fresh data for a fixed compute budget; after that, returns faded Established. When even that is not enough, models can write data for other models. How far synthetic data can replace human writing is an open question Active research.
Scraping the web has a catch for evaluation: benchmarks live on the web too. GPT-3's authors tried to remove benchmark test data from their training set, but a bug left some in and retraining was too expensive Established. This is Chapter 3's data leakage at the scale of the internet.
The interface
Building the tokenizer
Chapter 8 trained a toy BPE tokenizer. A real one is trained the same way on a large sample of the training mix, and then frozen: the model's embedding table is built around it, so it can't change later without retraining.
The main dial is vocabulary size. A larger vocabulary packs more text into each token, so the same compute and context window cover more text, at the cost of a bigger embedding table. Llama 3 grew its vocabulary to 128,000 tokens, which raised English compression from 3.17 to 3.94 characters per token compared with Llama 2 Established: about 19% fewer tokens for the same text, for the model's whole life. Other choices are rules (LLaMA split every number into single digits and fell back to raw bytes for unknown characters Established) and special tokens that mark structure, such as the end of a document or, later, the turns of a chat.
The engine
The training loop
The loop is Chapter 4's: forward pass, loss, backward pass, update. What changes is size. Batches are counted in tokens: Llama 3 405B started at 4 million tokens per step and doubled twice, to 16 million Established. No GPU holds that many at once, so each processes a few sequences at a time and accumulates gradients until the batch is complete, and only then does the optimizer step. With trillions of tokens, most data is seen about once, so classic overfitting rarely shows up and the training loss on fresh batches is itself an honest measure of progress.
The learning rate is not fixed. It warms up from near zero, because early steps on random weights with immature optimizer statistics can wreck a run, then decays along a cosine curve so the model can settle. Gradient clipping caps any single oversized step.
A linear warmup to 8 × 10⁻⁵ over the first 8,000 steps (dashed line), then a cosine decay to 8 × 10⁻⁷ by step 1.2 million. The warmup is under 1% of the run, too short to see on the left.
The arithmetic
Numbers in sixteen bits
GPUs multiply 16-bit numbers much faster than 32-bit ones, and 16-bit activations take half the memory. But 16 bits lose things. In FP16, gradients smaller than about 6 × 10⁻⁸ become zero, and with only a few significant digits, 1.0 + 0.0001 rounds back to 1.0, so small updates vanish.
Mixed precision keeps the best of both. The forward and backward passes run in 16 bits, and a 32-bit master copy of the weights receives the updates. With FP16, the loss is scaled up so small gradients survive. BF16, a 16-bit format with FP32's range but fewer digits, trains without loss scaling Established, which is why it is the usual choice today.
The first wall
Where the memory goes
The ZeRO paper starts from a puzzle: a 1.5-billion-parameter model needs only 3 GB for its 16-bit weights, yet could not be trained on a 32 GB GPU Established. The answer is the training state. With mixed-precision Adam, every parameter carries a 16-bit weight and gradient (2 + 2 bytes) plus a 32-bit master weight and Adam's two running averages (4 + 4 + 4): 16 bytes per parameter. A 7-billion-parameter model's weights fit on one 80 GB GPU; its training state, about 107 GB, does not.
Then come the activations, the intermediate values saved for the backward pass. They grow with sequence length, batch size and depth. For LLaMA 7B with one 4,096-token sequence, storing all of them takes about as much memory again as the model states. Recomputation keeps only some and recomputes the rest during the backward pass, trading extra compute for memory.
Start the calculator below at its defaults: 211 GB per GPU, hopeless. Shard the model states across 8 GPUs with ZeRO-3 and it is still about 118 GB, because sharding doesn't touch activations. Now recompute the attention scores: about 32 GB, and it fits.
Try it
Add up the memory a GPU needs to train a real model: weights, gradients, optimizer states and activations. Then shard and recompute until it fits in 80 GB.
The second wall
Splitting the work
Even when memory allows it, one GPU would take centuries. The work has to be spread over thousands, in up to four ways at once.
- Data parallelism: every GPU holds a copy of the model and processes a different slice of the batch, and an all-reduce averages their gradients before each update. The maths is unchanged; it is map-then-reduce.
- Sharding (ZeRO, FSDP): instead of identical copies, each GPU keeps only a slice of the weights, gradients and optimizer states, gathering what it needs just in time. In ZeRO's example, a 7.5B model's states fall from 120 GB per GPU to 1.9 GB on 64 GPUs Established.
- Tensor parallelism: split each layer's matrices across the GPUs of one server, which then compute every layer together, talking constantly over very fast links.
- Pipeline parallelism: give each server a block of consecutive layers and stream micro-batches through them like an assembly line, at the cost of an idle "bubble" while the pipeline fills and drains.
- Each GPU holds
- A full copy of the model, and a different slice of the batch.
- What they exchange
- Once per step, every GPU averages its gradients with all the others (an all-reduce), so all copies stay identical.
- Used
- Always, as the outermost layer of parallelism: it is how a run uses thousands of GPUs.
- The catch
- The whole model, its gradients and optimizer states must fit on every GPU.
Real runs combine them. Llama 3 405B used tensor parallelism across the 8 NVLink-connected GPUs of each server, 16 pipeline stages, and sharded data parallelism (FSDP) across the rest, plus a fourth kind, context parallelism, for very long sequences: up to 16,384 GPUs.
Llama 3 405B combined all of them: tensor parallelism across the 8 GPUs in each server, 16 pipeline stages, 128-way sharded data parallelism, and context parallelism for very long sequences, at 38–43% of the GPUs' peak arithmetic rate Established.
The long run
When things break
A synchronous job on thousands of GPUs, running for months, will break. During one 54-day period of Llama 3 405B's training there were 466 interruptions, 419 of them unexpected, and about 78% of those were attributed to hardware, mostly GPUs Established: one about every three hours. Automation still kept more than 90% of the time spent on useful training Established. The defence is frequent checkpoints, automatic restarts and tools that find the faulty machine quickly.
The optimisation can break too. PaLM's largest run saw about 20 loss spikes; the team restarted from about 100 steps before each spike and skipped a few hundred batches Established. Why such spikes happen in very large models is not fully understood Active research.
The bill
Counting the compute
There is a one-line estimate of the whole bill. Each parameter is used about once per token on the forward pass (a multiply and an add, 2 FLOPs) and about twice as much on the backward pass (4 FLOPs). So training N parameters on D tokens takes
floating-point operations. Check it: 6 × 405 × 10⁹ × 15.6 × 10¹² ≈ 3.8 × 10²⁵, exactly the figure Meta reports. For GPT-3, 6 × 175 × 10⁹ × 300 × 10⁹ ≈ 3.15 × 10²³, against a reported 3.14 × 10²³.
To turn FLOPs into time, divide by what the hardware actually sustains, not its peak. Llama 3 reached about 400 trillion FLOP/s per H100 Established. On 16,384 GPUs that gives 3.8 × 10²⁵ ÷ (16,384 × 4 × 10¹⁴) ≈ 5.8 million seconds: about 67 days of uninterrupted training. That is our estimate, before any time lost to failures.
The decision
Spending the budget
With C fixed, D = C / 6N. Double the model and you halve its tokens. Which split gives the best model? Kaplan and colleagues (2020) concluded that most extra compute should go into parameters Established, and models grew much faster than their datasets: GPT-3 saw under 2 tokens per parameter. Hoffmann and colleagues (2022) found instead that parameters and tokens should grow about equally, and showed it with Chinchilla: 70B parameters on 1.4 trillion tokens, the same compute as the 280B Gopher, and better Established. About 20 tokens per parameter became the rule of thumb.
The explorer below uses the scaling-law fit behind that result. It starts at Gopher's size for Gopher's budget. Slide the model size to the left and the predicted loss falls, then rises again: the bottom sits near 72B parameters, almost exactly Chinchilla. Then switch to the fit as printed in the 2022 paper. Its advice is quite different, and a 2024 replication traced the difference to the original fitting procedure and rounded coefficients Established. Influential numbers deserve checking.
Try it
Split a fixed number of FLOPs between model size and training tokens with a published scaling law, place real models on the map, and see how serving costs change the answer.
"Optimal" here means cheapest to train. A deployed model also costs about 2N FLOPs for every token it generates, possibly trillions of them. Move the last slider and the cheapest model for the same quality gets smaller and is trained for longer. Llama 3's flagship is roughly compute-optimal, but its smaller models were trained far longer than compute-optimal, and performed better than compute-optimal models at the same inference cost Established.
Why it matters
Why it matters
Chapter 8 showed that next-token prediction at scale works. This chapter shows that "at scale" is mostly engineering: data pipelines, numerical formats, memory accounting, distributed systems, reliability and budgeting, each with its own trade-offs.
- For an engineer, the numbers here are napkin maths you can reuse: 16 bytes per parameter to train, 2 to run, 6ND FLOPs to train, about 2N per generated token, 40% of peak throughput. They tell you in a minute whether a fine-tuning job fits on your GPUs and roughly what it costs.
- For a researcher, the open questions are just as concrete: which data actually helps, how far synthetic data can go, why large runs spike, how far scaling laws extrapolate, and how to evaluate models trained on the whole web.
What comes next
What comes out
At the end of all this, months of compute and millions of dollars, you have a base model. It has read trillions of tokens and predicts the next one extremely well. But ask it a question and it may continue with three more questions, as a quiz page would. Chapter 10 covers how post-training (supervised fine-tuning and preference optimisation) turns this raw text predictor into an assistant.
From a web crawl to a trained base model
- Building a Pretraining Dataset
- Deduplication and MinHash
- Data Mixtures, Repetition and Synthetic Data
- Designing the Tokenizer
- The Pretraining Loop
- Mixed-Precision Training (FP16 and BF16)
- Where Training Memory Goes
- Data Parallelism and All-Reduce
- The Compute Budget (C ≈ 6ND)
- Compute-Optimal Training (Kaplan vs Chinchilla)
- Post-training (Chapter 10)
Concepts in this chapter
Mark each one as you go. Must-know concepts are the core path.
- The Compute Budget (C ≈ 6ND)Training a dense Transformer with N parameters on D tokens costs about 6ND floating-point operations, which, divided by what a GPU actually sustains, turns a model plan into GPU-days and money.Know wellMust know
- Compute-Optimal Training (Kaplan vs Chinchilla)For a fixed compute budget, the lowest loss comes from growing parameters and training tokens together, about 20 tokens per parameter by the Chinchilla analysis, but labs deliberately train smaller models on far more tokens when the model will serve many users.Know wellMust know
- Data Mixtures, Repetition and Synthetic DataA training set is a weighted mixture of sources (web, code, books, papers, maths), and the weights, the number of times each source is repeated, and any model-generated data are deliberate design choices.UnderstandMust know
- Data Parallelism and All-ReduceData parallelism runs a copy of the model on every GPU, gives each a different slice of the batch, and averages their gradients with an all-reduce before every update, so all copies stay identical.Know wellMust know
- Deduplication and MinHashDeduplication removes repeated and near-identical documents from training data, usually by comparing sets of overlapping word sequences with MinHash signatures, so that copies are neither memorised nor overweighted.Know wellMust know
- Learning-Rate Schedules, Warmup and ClippingThe learning rate is ramped up from near zero during a short warmup, then decayed (usually along a cosine curve) to a small fraction of its peak, while gradient clipping caps any single step that would be too large.Know wellMust know
- Mixed-Precision Training (FP16 and BF16)Mixed-precision training does most arithmetic in 16-bit floating point, which is faster and halves activation memory, while keeping a 32-bit master copy of the weights so tiny updates are not lost.UnderstandMust know
- Tensor and Pipeline ParallelismModel parallelism splits the model itself: tensor parallelism divides each layer's matrices across GPUs, pipeline parallelism gives each GPU a block of consecutive layers, and large runs combine both with data parallelism.UnderstandMust know
- Building a Pretraining DatasetA pretraining dataset is built by a pipeline that extracts text from web crawls and other sources, then removes most of it, through language identification, quality filters, safety filters and deduplication, before trillions of tokens remain.Know wellMust know
- The Pretraining LoopPretraining repeats one step about a million times: take a batch of millions of tokens, compute the average next-token loss, backpropagate, and let the optimizer nudge every weight, usually seeing each piece of data about once.Know wellMust know
- Where Training Memory GoesTraining needs memory for the weights, their gradients, the optimizer's states (about 16 bytes per parameter in total with mixed-precision Adam) and the activations saved for the backward pass, which grow with sequence length and batch size.Know wellMust know
- Benchmark ContaminationContamination happens when benchmark test questions (or their answers) end up in a model's training data, so its score partly measures memory rather than ability.UnderstandShould know
- Sharded Training (ZeRO and FSDP)Sharded training keeps data parallelism's simplicity but stores only a slice of the optimizer states, gradients and weights on each GPU, gathering each layer's full weights just in time to use them.UnderstandShould know
- Designing the TokenizerBuilding a tokenizer means choosing an algorithm (BPE, WordPiece or unigram), training data, a vocabulary size, rules for numbers and unknown characters, and a set of special tokens, and every choice is frozen into the model for good.Know wellShould know
- Loss Spikes, Failures and CheckpointsLong training runs break: the loss can suddenly spike or diverge, and among thousands of GPUs something fails every few hours, so runs save checkpoints often and are built to roll back and resume.UnderstandShould know
What do I actually need to remember?
- Most of a web crawl is thrown away: text extraction, language ID, rule and model-based quality filters, safety filters and deduplication leave trillions of tokens from far more.
- Near-duplicates are found by comparing shingle sets with MinHash signatures and banded hashing; copies are memorised and leak into test sets.
- A training set is a weighted mixture; valuable sources are repeated (up to about 4 epochs is nearly as good as new data) and the weights are tuned with small models.
- The tokenizer is designed and frozen before training: vocabulary size trades compression against embedding size, and special tokens mark structure.
- A step averages the loss over millions of tokens; gradients are accumulated over micro-batches; learning rates warm up, then decay along a cosine.
- Mixed precision computes in 16 bits (usually BF16) and keeps FP32 master weights; with Adam, model states still cost about 16 bytes per parameter.
- Training memory = model states (16 bytes per parameter) + activations, which grow with sequence length and batch; recomputation trades compute for memory.
- Data parallelism copies the model and averages gradients; ZeRO/FSDP shards the copies; tensor and pipeline parallelism split the model itself.
- Training compute ≈ 6 × parameters × tokens; divide by sustained GPU throughput (about 40% of peak) to get GPU-days.
- For a fixed budget, grow parameters and tokens together (about 20 tokens per parameter); models meant for heavy use are trained smaller and longer.
You do not need to memorize everything else. This list is the revision sheet.
Key papers
The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale
Guilherme Penedo, Hynek Kydlíček et al. · 2024 · NeurIPS 2024
The most thoroughly documented open recipe for turning Common Crawl into pretraining data, with an ablation for every step.
- Problem
- The datasets behind the best open models were not released, and little was known about which cleaning steps actually help.
- What was new
- A 15-trillion-token dataset from 96 Common Crawl snapshots, per-snapshot MinHash deduplication (global deduplication worked worse), and FineWeb-Edu: 1.3T tokens chosen by a classifier trained on Llama 3's ratings of educational value.
How to read it: Section 3 walks through the pipeline step by step; the global-deduplication surprise is in 3.4.
CCNet: Extracting High Quality Monolingual Datasets from Web Crawl Data
Guillaume Wenzek, Marie-Anne Lachaux et al. · 2019
A widely copied web-cleaning pipeline: deduplicate, identify the language, then keep documents that a language model trained on Wikipedia finds plausible.
- Problem
- Raw Common Crawl text is mostly boilerplate, duplicates and low-quality pages, in hundreds of languages.
- What was new
- Paragraph-level deduplication, fastText language identification, and quality filtering by perplexity under a small language model trained on high-quality text.
Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus
Jesse Dodge, Maarten Sap et al. · 2021
One of the first audits of what a web-scale training corpus actually contains, including benchmark test data and the side effects of its filters.
- Problem
- Pretraining corpora were released with almost no documentation of their contents.
- What was new
- Found patents, machine-generated text and examples from NLP benchmarks inside C4, and showed that its blocklist filter disproportionately removed text from and about minority groups.
On the resemblance and containment of documents
Andrei Z. Broder · 1997 · Compression and Complexity of SEQUENCES 1997
Introduced the shingling and min-wise hashing (MinHash) technique that pretraining pipelines still use to find near-duplicate web pages at scale.
- Problem
- Comparing every pair of documents in a web-scale collection is far too slow, and exact hashing misses copies with small edits.
- What was new
- Represent a document by its set of overlapping word sequences (shingles), and estimate the overlap between two sets from small fixed-size sketches of minimum hash values.
How to read it: Read the definitions of resemblance and the sketching argument; skip the containment variant on a first pass.
Deduplicating Training Data Makes Language Models Better
Katherine Lee, Daphne Ippolito et al. · 2021
Showed that standard training sets are full of duplicates, and that removing them reduces memorisation and leaks between training and test data.
- Problem
- Repeated text in the training data gets memorised and overweighted, and copies of test examples inflate evaluation scores.
- What was new
- Exact-substring and MinHash deduplication at scale. One 61-word sentence appeared over 60,000 times in C4; deduplicated models emitted memorised text ten times less often.
The Pile: An 800GB Dataset of Diverse Text for Language Modeling
Leo Gao, Stella Biderman et al. · 2020
An influential open pretraining dataset that mixed 22 sources, from academic papers to code, and documented them.
- Problem
- Models trained mostly on web text lacked knowledge from specialised, high-quality domains, and few training sets were open.
- What was new
- An 825 GiB English corpus assembled from 22 diverse subsets, released with documentation and analysis.
LLaMA: Open and Efficient Foundation Language Models
Hugo Touvron, Thibaut Lavril et al. · 2023
Showed that smaller models trained on more tokens, using only publicly available data, can rival much larger ones, and released weights to researchers, starting the open-weight wave.
- Problem
- The strongest language models were closed, and very large.
- What was new
- Models from 7B to 65B parameters trained on trillions of tokens of public data; the 13B model outperformed GPT-3 (175B) on most benchmarks reported.
Scaling Data-Constrained Language Models
Niklas Muennighoff, Alexander M. Rush et al. · 2023
Asked what happens when high-quality text runs out: how much is repeated data worth?
- Problem
- Scaling laws assume fresh data, but the supply of good text is finite.
- What was new
- Up to about 4 epochs, repeated data was almost as good as new data for a fixed compute budget; the value of further repetition decayed, approaching zero after many more epochs.
Textbooks Are All You Need
Suriya Gunasekar, Yi Zhang et al. · 2023
A striking demonstration that carefully selected and synthetic data can let a small model compete with much larger ones on a narrow task.
- Problem
- Most code on the web is poor teaching material; does data quality matter as much as quantity?
- What was new
- phi-1, a 1.3B model trained on 6B tokens of filtered 'textbook quality' web code plus 1B tokens of textbooks and exercises generated by GPT-3.5, reached 50.6% pass@1 on HumanEval.
How to read it: Read it as evidence about data quality, not as a general recipe: the evaluation is narrow (Python functions).
Japanese and Korean voice search
Mike Schuster, Kaisuke Nakajima · 2012 · ICASSP 2012
The origin of the WordPiece subword method, later used for BERT's 30,000-token vocabulary.
- Problem
- Japanese and Korean text has huge character inventories and no spaces between words, so word-level vocabularies fail.
- What was new
- Build a vocabulary of word pieces greedily, adding the unit that most increases the likelihood of the training data under a language model.
How to read it: Only the section on building the word inventory matters here; the rest is about speech recognition.
SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing
Taku Kudo, John Richardson · 2018
The tokenizer library behind many LLMs, including LLaMA: it trains BPE or unigram vocabularies directly from raw text in any language.
- Problem
- Subword tools assumed text already split into words by spaces, which fails for Chinese, Japanese and similar languages.
- What was new
- Treat the input as a raw character stream, encoding spaces as an ordinary symbol, so tokenization is fully reversible and language independent.
SGDR: Stochastic Gradient Descent with Warm Restarts
Ilya Loshchilov, Frank Hutter · 2016 · ICLR 2017
Introduced the cosine learning-rate schedule that most LLM pretraining runs still use, without the restarts.
- Problem
- Step-wise learning-rate drops are abrupt and need hand-tuned timing.
- What was new
- Decay the learning rate smoothly along half a cosine curve, optionally restarting it periodically.
How to read it: Equation 5 is the whole idea.
Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
Priya Goyal, Piotr Dollár et al. · 2017
Showed how to train with very large batches across many GPUs without losing accuracy, using a linear learning-rate scaling rule and a gradual warmup.
- Problem
- Spreading training over many GPUs means large batches, and large batches caused optimisation trouble early in training.
- What was new
- Scale the learning rate with the batch size and ramp it up over the first few epochs: ResNet-50 trained with 8,192-image batches on 256 GPUs in one hour.
Decoupled Weight Decay Regularization
Ilya Loshchilov, Frank Hutter · 2017 · ICLR 2019
Introduced AdamW, the variant of Adam used to train most modern Transformers and LLMs.
- Problem
- With Adam, the usual L2 regularization doesn't behave like true weight decay, hurting generalization.
- What was new
- Apply weight decay directly to the weights, separately ('decoupled') from Adam's adaptive gradient step.
Mixed Precision Training
Paulius Micikevicius, Sharan Narang et al. · 2017 · ICLR 2018
The recipe for training in 16-bit floating point without losing accuracy, which roughly halves activation memory and unlocks GPUs' fastest arithmetic.
- Problem
- FP16's narrow range makes small gradients round to zero and weight updates vanish.
- What was new
- Keep an FP32 master copy of the weights, scale the loss up so small gradients survive in FP16, and accumulate products in FP32.
A Study of BFLOAT16 for Deep Learning Training
Dhiraj Kalamkar, Dheevatsa Mudigere et al. · 2019
Showed that the bfloat16 format, with FP32's range and fewer digits of precision, trains networks to the same accuracy without loss scaling.
- Problem
- FP16 training needs loss scaling and careful tuning because its range is so narrow.
- What was new
- A broad study across vision, speech, language and recommendation models showing BF16 matches FP32 accuracy with no hyperparameter changes.
- Built on
- Mixed Precision Training
ZeRO: Memory Optimizations Toward Training Trillion Parameter Models
Samyam Rajbhandari, Jeff Rasley et al. · 2019
Explained where training memory goes (16 bytes per parameter with mixed-precision Adam) and how to remove the redundant copies. Its stages became DeepSpeed's ZeRO and PyTorch's FSDP.
- Problem
- Data parallelism keeps a full copy of the weights, gradients and optimizer states on every GPU, so models that don't fit on one GPU can't use it.
- What was new
- Shard the optimizer states, then the gradients, then the parameters across data-parallel GPUs, gathering each piece only when needed: memory per GPU falls in proportion to the number of GPUs.
How to read it: Section 3 ('Where did all the memory go?') and Figure 1 are the essentials; the rest is engineering detail.
Training Deep Nets with Sublinear Memory Cost
Tianqi Chen, Bing Xu et al. · 2016
Activation checkpointing: trade a little extra computation for a large cut in training memory. Every large model run uses some form of it.
- Problem
- Storing every intermediate activation for the backward pass makes memory grow linearly with depth.
- What was new
- Store activations only at checkpoints and recompute the rest during the backward pass: O(√n) memory for an n-layer network at the cost of about one extra forward pass.
Reducing Activation Recomputation in Large Transformer Models
Vijay Korthikanti, Jared Casper et al. · 2022
Worked out exactly how much activation memory a Transformer layer needs, and how to recompute only the cheap, memory-hungry parts.
- Problem
- Recomputing every layer saves memory but costs 30–40% more compute.
- What was new
- A per-layer activation formula, sequence parallelism, and selective recomputation of the attention scores, which cuts activation memory about fivefold with little overhead.
How to read it: Section 4.1 derives the 34 + 5as/h formula used in this chapter's memory calculator.
Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Mohammad Shoeybi, Mostofa Patwary et al. · 2019
Introduced the tensor-parallel layout for Transformers that most large training systems still use: split each layer's matrices across GPUs.
- Problem
- Multi-billion-parameter Transformers no longer fit on a single GPU.
- What was new
- Split the attention heads and the feed-forward matrices across GPUs so each layer needs only a couple of all-reduces; trained an 8.3B model on 512 GPUs at 76% scaling efficiency.
- Built on
- Attention Is All You Need
How to read it: Figure 3 (how an MLP and an attention block are split) is the part to understand.
GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism
Yanping Huang, Youlong Cheng et al. · 2018
Made pipeline parallelism practical: split a network's layers across accelerators and keep them busy by streaming micro-batches through.
- Problem
- A model too deep for one accelerator can be split by layers, but then only one accelerator works at a time.
- What was new
- Split each batch into micro-batches and pipeline them through the stages; the idle 'bubble' becomes negligible once there are at least four times as many micro-batches as stages.
Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Deepak Narayanan, Mohammad Shoeybi et al. · 2021
Showed how to compose tensor, pipeline and data parallelism to train trillion-parameter models efficiently on thousands of GPUs.
- Problem
- Each form of parallelism stops scaling on its own: tensor parallelism across servers is too slow, and pipelines waste time in bubbles.
- What was new
- Practical rules for combining the three (tensor parallelism inside a server, pipelines across servers) and an interleaved schedule: a 1-trillion-parameter model at 502 petaFLOP/s on 3,072 GPUs, 52% of peak.
How to read it: Read the takeaways in Section 3; they summarise how to choose the parallel sizes.
PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel
Yanli Zhao, Andrew Gu et al. · 2023 · PVLDB 16(12)
Describes PyTorch's built-in version of ZeRO-style sharding, the default way to train models that don't fit on one GPU.
- Problem
- ZeRO's ideas needed a general, efficient implementation inside the main deep-learning framework.
- What was new
- Fully sharded data parallelism in PyTorch, with the engineering for overlapping communication and computation and for mixed precision.
PaLM: Scaling Language Modeling with Pathways
Aakanksha Chowdhery, Sharan Narang et al. · 2022
A 540-billion-parameter model whose report defined model FLOPs utilization (MFU) and described, unusually frankly, the loss spikes of a very large run.
- Problem
- How efficiently can a dense model be trained across thousands of accelerators, and how do you keep a giant run stable?
- What was new
- Training across two TPU v4 pods at 46.2% MFU; loss spikes (about 20 in the largest run) handled by restarting from an earlier checkpoint and skipping the surrounding batches.
How to read it: Sections 4 (training infrastructure) and 5.1 (training instability) are the parts for this chapter.
OPT: Open Pre-trained Transformer Language Models
Susan Zhang, Stephen Roller et al. · 2022
Released a GPT-3-sized model with its full training logbook: a rare, honest record of what goes wrong in a large run.
- Problem
- How large models are actually trained, including failures and mid-run fixes, was rarely documented.
- What was new
- OPT-175B on 992 A100 GPUs, with at least 35 manual restarts from hardware failures over two months, and loss divergences handled by lowering the learning rate and restarting from earlier checkpoints.
How to read it: Section 2.5, 'Training Processes', and the released logbook.
Scaling Laws for Neural Language Models
Jared Kaplan, Sam McCandlish et al. · 2020
Found that language-model loss falls as a smooth power law in parameters, data and compute — making model scale something you could plan.
- Problem
- There was no quantitative way to predict how much better a larger model would be.
- What was new
- Empirical power-law fits of loss against model size, dataset size and compute, over many orders of magnitude.
- Built on
- Attention Is All You Need
Training Compute-Optimal Large Language Models
Jordan Hoffmann, Sebastian Borgeaud et al. · 2022 · NeurIPS 2022
Showed that many large models were undertrained: for a fixed compute budget, parameters and training tokens should grow roughly in proportion.
- Problem
- Earlier scaling recommendations favoured very large models trained on comparatively little data.
- What was new
- Trained 400+ models to fit compute-optimal trade-offs; the 70B 'Chinchilla' model outperformed much larger models trained on fewer tokens.
Chinchilla Scaling: A replication attempt
Tamay Besiroglu, Ege Erdil et al. · 2024
Re-fitted Chinchilla's scaling law from its published data and found the printed coefficients inconsistent with the paper's own conclusions: a lesson in checking influential numbers.
- Problem
- Chinchilla's third method gave a fit that recommended far more tokens per parameter than its other two methods, with implausibly tight confidence intervals.
- What was new
- A re-fit (E = 1.82, A = 482, B = 2085, α = 0.348, β = 0.366) consistent with about 20 tokens per parameter, and an explanation: the original optimiser stopped early and the printed values were rounded.
How to read it: Short and readable; a good model of how to replicate a result from a figure.
Beyond Chinchilla-Optimal: Accounting for Inference in Language Model Scaling Laws
Nikhil Sardana, Jacob Portes et al. · 2023
Formalised why labs train models far past the 'compute-optimal' point: the model will be run billions of times.
- Problem
- Chinchilla's rule minimises training cost only, but a deployed model's inference cost can dwarf it.
- What was new
- Modified the Chinchilla laws to include inference demand; with large expected demand (around a billion requests), the cheapest model is smaller and trained on more data.
The Llama 3 Herd of Models
Aaron Grattafiori, Abhimanyu Dubey et al. · 2024
The most complete public account of building a frontier-scale model end to end: data pipeline, scaling-law experiments, 16,384-GPU training, failures and all.
- Problem
- Frontier labs had largely stopped publishing how their models were built.
- What was new
- A 405B model trained on 15.6T tokens with 3.8 × 10²⁵ FLOPs, with detailed data cleaning, a data mix chosen by scaling experiments, 4D parallelism at 38–43% utilisation, and 466 job interruptions in a 54-day period.
How to read it: It is 90+ pages. For this chapter read Section 3 (pre-training) only: data, scaling laws, infrastructure and the training recipe.
Watch
Stanford Online
Stanford CS336 Language Modeling from Scratch | Spring 2025 | Lecture 13: Data 1
A tour of what real pretraining datasets contain and how they were built, from a course that argues data matters most.
Covers: Pre-, mid- and post-training data, Common Crawl, C4 and other datasets, and the conversion, filtering and deduplication steps between a raw crawl and training text.
Stanford Online
Stanford CS336 Language Modeling from Scratch | Spring 2025 | Lec. 2: Pytorch, Resource Accounting
Teaches the habit this chapter is built on: napkin maths for memory (bytes per parameter) and compute (FLOPs) before you train anything.
Covers: Floating-point formats and their memory cost, counting the FLOPs of a training step, and the training loop from tensors to optimizer.
Stanford Online
Stanford CS336 Language Modeling from Scratch | Spring 2025 | Lecture 7: Parallelism 1
A clear university lecture on how training is split across many GPUs, with the trade-offs between the methods.
Covers: Data parallelism, ZeRO stages 1–3 and FSDP, tensor and pipeline parallelism, activation recomputation, and how they combine.
Stanford Online
Stanford CS336 Language Modeling from Scratch | Spring 2025 | Lecture 9: Scaling laws 1
Goes beyond 'bigger is better' to how scaling laws are fitted and used to make decisions.
Covers: Data scaling, Kaplan et al. and Chinchilla, IsoFLOP analysis and choosing the batch size.
Andrej Karpathy
Let's reproduce GPT-2 (124M)
Build and train the smallest GPT-2 from scratch in PyTorch, then optimise it. For when you want to implement what this chapter describes.
Covers: The GPT-2 network in code, the engineering that makes training fast, and a full training run using the GPT-2 and GPT-3 papers' hyperparameters.
Andrej Karpathy
Deep Dive into LLMs like ChatGPT
A long, general-audience walk through the whole pipeline behind a chat model, from internet text to tokens to pretraining to post-training.
Covers: The full training stack behind a chat model, from internet text and tokens to a pretrained base model and the post-training that turns it into an assistant.
What came next?
Chapter 10
From Base Model to Assistant →
A base model can answer questions, but pretraining alone does not make it reliably follow instructions or conversation roles.