Skip to content
Road to Intelligence

Part III · Large Language Models

Chapter 9

How an LLM Is Actually Built

Data, tokens, GPUs and the laws of scale.

2 h 30 min core path15 concepts3 interactivesCore path · 4 optional concepts foldedDeep · all 15 concepts shown in full

In one sentenceBuilding an LLM is a data, systems and budgeting problem: curate trillions of tokens, split training across thousands of accelerators, and spend compute where scaling laws say it pays.

The problem

The shopping list

Chapter 8 summed up how large language models are made in one sentence: predict the next token, at scale. This chapter is about the words "at scale". What would it actually take to train one?

Most frontier labs no longer say. One exception is Meta's report on Llama 3, which describes training a 405-billion-parameter model in unusual detail, failures included. It makes a good running example. Here is its bill of materials:

WhatLlama 3 405BSection
Training text15.6 trillion tokens, cleaned from far moreTrillions of tokens; The mix
Tokenizer128,000-token vocabularyThe tokenizer
Training stepsabout 1.2 million, each over millions of tokensThe training loop
Memory for training stateabout 6.5 TB (16 bytes per parameter)Memory
Hardwareup to 16,384 H100 GPUs, 80 GB eachMany GPUs
Interruptions466 in one 54-day periodWhen things break
Compute3.8 × 10²⁵ floating-point operationsCounting the compute

The token count, vocabulary, GPU count, interruptions and compute are all from the Llama 3 report Established. The 6.5 TB is our arithmetic: 405 billion parameters at 16 bytes each, a figure explained below.

Read down the list and it is less a machine-learning problem than a data-engineering and distributed-systems one. If you build data pipelines for a living, you already know half of this chapter: ETL, deduplication, sharding, checkpoints and capacity planning. The other half is what those words mean when the job is a neural network.

The raw material

Trillions of tokens

The text mostly comes from Common Crawl, a public archive of the web released as snapshots. Raw, it is mostly menus, cookie banners, product listings, spam and copies. A pretraining pipeline extracts the main text from each page, identifies its language, applies cheap rules (too short, too repetitive, too many symbols), scores quality with a classifier, removes unsafe and personal content, and deduplicates.

Most of what goes in does not come out. In the open FineWeb dataset, basic filtering of 96 Common Crawl snapshots left about 36 trillion tokens; deduplication left about 20 trillion; further filters left 15 trillion. An educational-quality classifier, trained on Llama 3's ratings of 460,000 pages, then selected 1.3 trillion of those for FineWeb-Edu Established.

Copies deserve their own step. One 61-word sentence appeared over 60,000 times in the C4 dataset, and models trained after deduplication emitted memorised text ten times less often Established. Exact hashing only catches identical pages, so pipelines compare documents as sets of overlapping word sequences (shingles) and estimate their overlap with compact MinHash signatures. Try the method on eleven made-up web pages. At the default settings, which are FineWeb's, the mirror and the syndicated copy of a news story are caught, as are two shop pages built from one template. A re-posted recipe with a reworded first sentence slips through. Lower the rows per band to catch it.

Try it · toy model

Find the Near-Duplicates

Run the deduplication method web-scale datasets use (shingles, MinHash and banded hashing) on a dozen made-up web pages, and tune what counts as a copy.

Know well8 min

Even this has surprises. FineWeb's team found that deduplicating all 96 snapshots against each other made the data worse than deduplicating each snapshot on its own: in the oldest crawls, what survived global deduplication was lower quality than what it removed Established.

The recipe

What goes in the mix

A training set is not one source but a weighted blend. Sampling in proportion to size would bury the small, dense sources (books, encyclopedias, code, papers) under web text. LLaMA (2023) drew 67% of its tokens from Common Crawl and 4.5% from Wikipedia, but went over Wikipedia about 2.45 times and over Common Crawl about 1.1 times Established. Llama 3 chose its mix by training small models on candidate blends and extrapolating with scaling laws, ending at roughly 50% general knowledge, 25% maths and reasoning, 17% code and 8% multilingual text Established.

How often can the good data be repeated? Up to about 4 epochs, repeated data was almost as valuable as fresh data for a fixed compute budget; after that, returns faded Established. When even that is not enough, models can write data for other models. How far synthetic data can replace human writing is an open question Active research.

Scraping the web has a catch for evaluation: benchmarks live on the web too. GPT-3's authors tried to remove benchmark test data from their training set, but a bug left some in and retraining was too expensive Established. This is Chapter 3's data leakage at the scale of the internet.

OptionalBenchmark Contamination· folded on the core path. Open it, or switch to Deep to show it here.

The interface

Building the tokenizer

Chapter 8 trained a toy BPE tokenizer. A real one is trained the same way on a large sample of the training mix, and then frozen: the model's embedding table is built around it, so it can't change later without retraining.

The main dial is vocabulary size. A larger vocabulary packs more text into each token, so the same compute and context window cover more text, at the cost of a bigger embedding table. Llama 3 grew its vocabulary to 128,000 tokens, which raised English compression from 3.17 to 3.94 characters per token compared with Llama 2 Established: about 19% fewer tokens for the same text, for the model's whole life. Other choices are rules (LLaMA split every number into single digits and fell back to raw bytes for unknown characters Established) and special tokens that mark structure, such as the end of a document or, later, the turns of a chat.

OptionalDesigning the Tokenizer· folded on the core path. Open it, or switch to Deep to show it here.

The engine

The training loop

The loop is Chapter 4's: forward pass, loss, backward pass, update. What changes is size. Batches are counted in tokens: Llama 3 405B started at 4 million tokens per step and doubled twice, to 16 million Established. No GPU holds that many at once, so each processes a few sequences at a time and accumulates gradients until the batch is complete, and only then does the optimizer step. With trillions of tokens, most data is seen about once, so classic overfitting rarely shows up and the training loss on fresh batches is itself an honest measure of progress.

The learning rate is not fixed. It warms up from near zero, because early steps on random weights with immature optimizer statistics can wreck a run, then decays along a cosine curve so the model can settle. Gradient clipping caps any single oversized step.

Learning rate over one real run · Llama 3 405B
0300k600k900k1200k04e-58e-5the whole run (steps)08k20k40k04e-58e-5the first 40,000 steps

A linear warmup to 8 × 10⁻⁵ over the first 8,000 steps (dashed line), then a cosine decay to 8 × 10⁻⁷ by step 1.2 million. The warmup is under 1% of the run, too short to see on the left.

The arithmetic

Numbers in sixteen bits

GPUs multiply 16-bit numbers much faster than 32-bit ones, and 16-bit activations take half the memory. But 16 bits lose things. In FP16, gradients smaller than about 6 × 10⁻⁸ become zero, and with only a few significant digits, 1.0 + 0.0001 rounds back to 1.0, so small updates vanish.

Mixed precision keeps the best of both. The forward and backward passes run in 16 bits, and a 32-bit master copy of the weights receives the updates. With FP16, the loss is scaled up so small gradients survive. BF16, a 16-bit format with FP32's range but fewer digits, trains without loss scaling Established, which is why it is the usual choice today.

The first wall

Where the memory goes

The ZeRO paper starts from a puzzle: a 1.5-billion-parameter model needs only 3 GB for its 16-bit weights, yet could not be trained on a 32 GB GPU Established. The answer is the training state. With mixed-precision Adam, every parameter carries a 16-bit weight and gradient (2 + 2 bytes) plus a 32-bit master weight and Adam's two running averages (4 + 4 + 4): 16 bytes per parameter. A 7-billion-parameter model's weights fit on one 80 GB GPU; its training state, about 107 GB, does not.

Then come the activations, the intermediate values saved for the backward pass. They grow with sequence length, batch size and depth. For LLaMA 7B with one 4,096-token sequence, storing all of them takes about as much memory again as the model states. Recomputation keeps only some and recomputes the rest during the backward pass, trading extra compute for memory.

Start the calculator below at its defaults: 211 GB per GPU, hopeless. Shard the model states across 8 GPUs with ZeRO-3 and it is still about 118 GB, because sharding doesn't touch activations. Now recompute the attention scores: about 32 GB, and it fits.

Try it

Will It Fit?

Add up the memory a GPU needs to train a real model: weights, gradients, optimizer states and activations. Then shard and recompute until it fits in 80 GB.

Know well8 min

The second wall

Splitting the work

Even when memory allows it, one GPU would take centuries. The work has to be spread over thousands, in up to four ways at once.

  • Data parallelism: every GPU holds a copy of the model and processes a different slice of the batch, and an all-reduce averages their gradients before each update. The maths is unchanged; it is map-then-reduce.
  • Sharding (ZeRO, FSDP): instead of identical copies, each GPU keeps only a slice of the weights, gradients and optimizer states, gathering what it needs just in time. In ZeRO's example, a 7.5B model's states fall from 120 GB per GPU to 1.9 GB on 64 GPUs Established.
  • Tensor parallelism: split each layer's matrices across the GPUs of one server, which then compute every layer together, talking constantly over very fast links.
  • Pipeline parallelism: give each server a block of consecutive layers and stream micro-batches through them like an assembly line, at the cost of an idle "bubble" while the pipeline fills and drains.
Four GPUs, one four-layer model · what each GPU holds
GPU 1batch slice 1L1L2L3L4GPU 2batch slice 2L1L2L3L4GPU 3batch slice 3L1L2L3L4GPU 4batch slice 4L1L2L3L4all-reduce gradients once per step
Each GPU holds
A full copy of the model, and a different slice of the batch.
What they exchange
Once per step, every GPU averages its gradients with all the others (an all-reduce), so all copies stay identical.
Used
Always, as the outermost layer of parallelism: it is how a run uses thousands of GPUs.
The catch
The whole model, its gradients and optimizer states must fit on every GPU.

Real runs combine them. Llama 3 405B used tensor parallelism across the 8 NVLink-connected GPUs of each server, 16 pipeline stages, and sharded data parallelism (FSDP) across the rest, plus a fourth kind, context parallelism, for very long sequences: up to 16,384 GPUs.

Llama 3 405B combined all of them: tensor parallelism across the 8 GPUs in each server, 16 pipeline stages, 128-way sharded data parallelism, and context parallelism for very long sequences, at 38–43% of the GPUs' peak arithmetic rate Established.

OptionalSharded Training (ZeRO and FSDP)· folded on the core path. Open it, or switch to Deep to show it here.

The long run

When things break

A synchronous job on thousands of GPUs, running for months, will break. During one 54-day period of Llama 3 405B's training there were 466 interruptions, 419 of them unexpected, and about 78% of those were attributed to hardware, mostly GPUs Established: one about every three hours. Automation still kept more than 90% of the time spent on useful training Established. The defence is frequent checkpoints, automatic restarts and tools that find the faulty machine quickly.

The optimisation can break too. PaLM's largest run saw about 20 loss spikes; the team restarted from about 100 steps before each spike and skipped a few hundred batches Established. Why such spikes happen in very large models is not fully understood Active research.

OptionalLoss Spikes, Failures and Checkpoints· folded on the core path. Open it, or switch to Deep to show it here.

The bill

Counting the compute

There is a one-line estimate of the whole bill. Each parameter is used about once per token on the forward pass (a multiply and an add, 2 FLOPs) and about twice as much on the backward pass (4 FLOPs). So training N parameters on D tokens takes

C≈6 N DC \approx 6\,N\,D

floating-point operations. Check it: 6 × 405 × 10⁹ × 15.6 × 10¹² ≈ 3.8 × 10²⁵, exactly the figure Meta reports. For GPT-3, 6 × 175 × 10⁹ × 300 × 10⁹ ≈ 3.15 × 10²³, against a reported 3.14 × 10²³.

To turn FLOPs into time, divide by what the hardware actually sustains, not its peak. Llama 3 reached about 400 trillion FLOP/s per H100 Established. On 16,384 GPUs that gives 3.8 × 10²⁵ ÷ (16,384 × 4 × 10¹⁴) ≈ 5.8 million seconds: about 67 days of uninterrupted training. That is our estimate, before any time lost to failures.

The decision

Spending the budget

With C fixed, D = C / 6N. Double the model and you halve its tokens. Which split gives the best model? Kaplan and colleagues (2020) concluded that most extra compute should go into parameters Established, and models grew much faster than their datasets: GPT-3 saw under 2 tokens per parameter. Hoffmann and colleagues (2022) found instead that parameters and tokens should grow about equally, and showed it with Chinchilla: 70B parameters on 1.4 trillion tokens, the same compute as the 280B Gopher, and better Established. About 20 tokens per parameter became the rule of thumb.

The explorer below uses the scaling-law fit behind that result. It starts at Gopher's size for Gopher's budget. Slide the model size to the left and the predicted loss falls, then rises again: the bottom sits near 72B parameters, almost exactly Chinchilla. Then switch to the fit as printed in the 2022 paper. Its advice is quite different, and a 2024 replication traced the difference to the original fitting procedure and rounded coefficients Established. Influential numbers deserve checking.

Try it

Spend a Compute Budget

Split a fixed number of FLOPs between model size and training tokens with a published scaling law, place real models on the map, and see how serving costs change the answer.

Know well9 min

"Optimal" here means cheapest to train. A deployed model also costs about 2N FLOPs for every token it generates, possibly trillions of them. Move the last slider and the cheapest model for the same quality gets smaller and is trained for longer. Llama 3's flagship is roughly compute-optimal, but its smaller models were trained far longer than compute-optimal, and performed better than compute-optimal models at the same inference cost Established.

Building at scale, 2019–2024Open the full timeline →
Transformers and scaleLLMs, agents and reasoning (evolving)2019: GPT-2GPT-220192019: ZeRO removes redundant copiesZeRO removes redundant copies20192020: Scaling lawsScaling laws20202020: GPT-3 and in-context learningGPT-3 and in-context learning20202020: Summaries learned from human feedbackSummaries learned from human feedback20202020: Vision TransformerVision Transformer20202020: AlphaFold 2AlphaFold 220202021: CLIP: images meet languageCLIP: images meet language20212021: LoRA: low-rank fine-tuningLoRA: low-rank fine-tuning20212022: Chain-of-thought promptingChain-of-thought prompting20222022: InstructGPT and RLHFInstructGPT and RLHF20222022: Chinchilla: compute-optimal trainingChinchilla: compute-optimal training20222022: FlashAttentionFlashAttention20222022: Open text-to-image modelsOpen text-to-image models20222022: ChatGPTChatGPT20222022: Constitutional AIConstitutional AI20222023: GPT-4GPT-420232023: Direct Preference OptimizationDirect Preference Optimization20232023: Capable open-weight LLMsCapable open-weight LLMs20232023: PagedAttention and vLLMPagedAttention and vLLM20232024: Mixtral: an open mixture-of-experts modelMixtral: an open mixture-of-experts model20242024: The Llama 3 report opens up a frontier-scale runThe Llama 3 report opens up…20242024: EU AI Act enters into forceEU AI Act enters into force20242024: Reasoning models: OpenAI o1Reasoning models: OpenAI o120242024: Nobel Prizes for AI researchNobel Prizes for…20242024: An open post-training recipe: Tulu 3An open post-training…20242024: DeepSeek-V3DeepSeek-V32024

Why it matters

Why it matters

Chapter 8 showed that next-token prediction at scale works. This chapter shows that "at scale" is mostly engineering: data pipelines, numerical formats, memory accounting, distributed systems, reliability and budgeting, each with its own trade-offs.

  • For an engineer, the numbers here are napkin maths you can reuse: 16 bytes per parameter to train, 2 to run, 6ND FLOPs to train, about 2N per generated token, 40% of peak throughput. They tell you in a minute whether a fine-tuning job fits on your GPUs and roughly what it costs.
  • For a researcher, the open questions are just as concrete: which data actually helps, how far synthetic data can go, why large runs spike, how far scaling laws extrapolate, and how to evaluate models trained on the whole web.

What comes next

What comes out

At the end of all this, months of compute and millions of dollars, you have a base model. It has read trillions of tokens and predicts the next one extremely well. But ask it a question and it may continue with three more questions, as a quiz page would. Chapter 10 covers how post-training (supervised fine-tuning and preference optimisation) turns this raw text predictor into an assistant.

Concepts in this chapter

Mark each one as you go. Must-know concepts are the core path.

What do I actually need to remember?

  • Most of a web crawl is thrown away: text extraction, language ID, rule and model-based quality filters, safety filters and deduplication leave trillions of tokens from far more.
  • Near-duplicates are found by comparing shingle sets with MinHash signatures and banded hashing; copies are memorised and leak into test sets.
  • A training set is a weighted mixture; valuable sources are repeated (up to about 4 epochs is nearly as good as new data) and the weights are tuned with small models.
  • The tokenizer is designed and frozen before training: vocabulary size trades compression against embedding size, and special tokens mark structure.
  • A step averages the loss over millions of tokens; gradients are accumulated over micro-batches; learning rates warm up, then decay along a cosine.
  • Mixed precision computes in 16 bits (usually BF16) and keeps FP32 master weights; with Adam, model states still cost about 16 bytes per parameter.
  • Training memory = model states (16 bytes per parameter) + activations, which grow with sequence length and batch; recomputation trades compute for memory.
  • Data parallelism copies the model and averages gradients; ZeRO/FSDP shards the copies; tensor and pipeline parallelism split the model itself.
  • Training compute ≈ 6 × parameters × tokens; divide by sustained GPU throughput (about 40% of peak) to get GPU-days.
  • For a fixed budget, grow parameters and tokens together (about 20 tokens per parameter); models meant for heavy use are trained smaller and longer.

You do not need to memorize everything else. This list is the revision sheet.

Key papers

Important

The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale

Guilherme Penedo, Hynek Kydlíček et al. · 2024 · NeurIPS 2024

The most thoroughly documented open recipe for turning Common Crawl into pretraining data, with an ablation for every step.

Problem
The datasets behind the best open models were not released, and little was known about which cleaning steps actually help.
What was new
A 15-trillion-token dataset from 96 Common Crawl snapshots, per-snapshot MinHash deduplication (global deduplication worked worse), and FineWeb-Edu: 1.3T tokens chosen by a classifier trained on Llama 3's ratings of educational value.

How to read it: Section 3 walks through the pipeline step by step; the global-deduplication surprise is in 3.4.

~45 min readarXiv:2406.17557✓ verified 2026-10-04
Optional

CCNet: Extracting High Quality Monolingual Datasets from Web Crawl Data

Guillaume Wenzek, Marie-Anne Lachaux et al. · 2019

A widely copied web-cleaning pipeline: deduplicate, identify the language, then keep documents that a language model trained on Wikipedia finds plausible.

Problem
Raw Common Crawl text is mostly boilerplate, duplicates and low-quality pages, in hundreds of languages.
What was new
Paragraph-level deduplication, fastText language identification, and quality filtering by perplexity under a small language model trained on high-quality text.
~20 min readarXiv:1911.00359✓ verified 2026-10-04
Optional

Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus

Jesse Dodge, Maarten Sap et al. · 2021

One of the first audits of what a web-scale training corpus actually contains, including benchmark test data and the side effects of its filters.

Problem
Pretraining corpora were released with almost no documentation of their contents.
What was new
Found patents, machine-generated text and examples from NLP benchmarks inside C4, and showed that its blocklist filter disproportionately removed text from and about minority groups.
~25 min readarXiv:2104.08758✓ verified 2026-10-04
Optional

On the resemblance and containment of documents

Andrei Z. Broder · 1997 · Compression and Complexity of SEQUENCES 1997

Introduced the shingling and min-wise hashing (MinHash) technique that pretraining pipelines still use to find near-duplicate web pages at scale.

Problem
Comparing every pair of documents in a web-scale collection is far too slow, and exact hashing misses copies with small edits.
What was new
Represent a document by its set of overlapping word sequences (shingles), and estimate the overlap between two sets from small fixed-size sketches of minimum hash values.

How to read it: Read the definitions of resemblance and the sketching argument; skip the containment variant on a first pass.

~25 min readdoi:10.1109/SEQUEN.1997.666900✓ verified 2026-10-04
Important

Deduplicating Training Data Makes Language Models Better

Katherine Lee, Daphne Ippolito et al. · 2021

Showed that standard training sets are full of duplicates, and that removing them reduces memorisation and leaks between training and test data.

Problem
Repeated text in the training data gets memorised and overweighted, and copies of test examples inflate evaluation scores.
What was new
Exact-substring and MinHash deduplication at scale. One 61-word sentence appeared over 60,000 times in C4; deduplicated models emitted memorised text ten times less often.
~30 min readarXiv:2107.06499✓ verified 2026-10-04
Optional

The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Leo Gao, Stella Biderman et al. · 2020

An influential open pretraining dataset that mixed 22 sources, from academic papers to code, and documented them.

Problem
Models trained mostly on web text lacked knowledge from specialised, high-quality domains, and few training sets were open.
What was new
An 825 GiB English corpus assembled from 22 diverse subsets, released with documentation and analysis.
~20 min readarXiv:2101.00027✓ verified 2026-10-04
Important

LLaMA: Open and Efficient Foundation Language Models

Hugo Touvron, Thibaut Lavril et al. · 2023

Showed that smaller models trained on more tokens, using only publicly available data, can rival much larger ones, and released weights to researchers, starting the open-weight wave.

Problem
The strongest language models were closed, and very large.
What was new
Models from 7B to 65B parameters trained on trillions of tokens of public data; the 13B model outperformed GPT-3 (175B) on most benchmarks reported.
~30 min readarXiv:2302.13971✓ verified 2026-09-26
Important

Scaling Data-Constrained Language Models

Niklas Muennighoff, Alexander M. Rush et al. · 2023

Asked what happens when high-quality text runs out: how much is repeated data worth?

Problem
Scaling laws assume fresh data, but the supply of good text is finite.
What was new
Up to about 4 epochs, repeated data was almost as good as new data for a fixed compute budget; the value of further repetition decayed, approaching zero after many more epochs.
~40 min readarXiv:2305.16264✓ verified 2026-10-04
Optional

Textbooks Are All You Need

Suriya Gunasekar, Yi Zhang et al. · 2023

A striking demonstration that carefully selected and synthetic data can let a small model compete with much larger ones on a narrow task.

Problem
Most code on the web is poor teaching material; does data quality matter as much as quantity?
What was new
phi-1, a 1.3B model trained on 6B tokens of filtered 'textbook quality' web code plus 1B tokens of textbooks and exercises generated by GPT-3.5, reached 50.6% pass@1 on HumanEval.

How to read it: Read it as evidence about data quality, not as a general recipe: the evaluation is narrow (Python functions).

~25 min readarXiv:2306.11644✓ verified 2026-10-04
Optional

Japanese and Korean voice search

Mike Schuster, Kaisuke Nakajima · 2012 · ICASSP 2012

The origin of the WordPiece subword method, later used for BERT's 30,000-token vocabulary.

Problem
Japanese and Korean text has huge character inventories and no spaces between words, so word-level vocabularies fail.
What was new
Build a vocabulary of word pieces greedily, adding the unit that most increases the likelihood of the training data under a language model.

How to read it: Only the section on building the word inventory matters here; the rest is about speech recognition.

~20 min readdoi:10.1109/ICASSP.2012.6289079✓ verified 2026-10-04
Important

SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing

Taku Kudo, John Richardson · 2018

The tokenizer library behind many LLMs, including LLaMA: it trains BPE or unigram vocabularies directly from raw text in any language.

Problem
Subword tools assumed text already split into words by spaces, which fails for Chinese, Japanese and similar languages.
What was new
Treat the input as a raw character stream, encoding spaces as an ordinary symbol, so tokenization is fully reversible and language independent.
~15 min readarXiv:1808.06226✓ verified 2026-10-04
Optional

SGDR: Stochastic Gradient Descent with Warm Restarts

Ilya Loshchilov, Frank Hutter · 2016 · ICLR 2017

Introduced the cosine learning-rate schedule that most LLM pretraining runs still use, without the restarts.

Problem
Step-wise learning-rate drops are abrupt and need hand-tuned timing.
What was new
Decay the learning rate smoothly along half a cosine curve, optionally restarting it periodically.

How to read it: Equation 5 is the whole idea.

~20 min readarXiv:1608.03983✓ verified 2026-10-04
Optional

Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour

Priya Goyal, Piotr Dollár et al. · 2017

Showed how to train with very large batches across many GPUs without losing accuracy, using a linear learning-rate scaling rule and a gradual warmup.

Problem
Spreading training over many GPUs means large batches, and large batches caused optimisation trouble early in training.
What was new
Scale the learning rate with the batch size and ramp it up over the first few epochs: ResNet-50 trained with 8,192-image batches on 256 GPUs in one hour.
~25 min readarXiv:1706.02677✓ verified 2026-10-04
Important

Decoupled Weight Decay Regularization

Ilya Loshchilov, Frank Hutter · 2017 · ICLR 2019

Introduced AdamW, the variant of Adam used to train most modern Transformers and LLMs.

Problem
With Adam, the usual L2 regularization doesn't behave like true weight decay, hurting generalization.
What was new
Apply weight decay directly to the weights, separately ('decoupled') from Adam's adaptive gradient step.
~40 min readarXiv:1711.05101✓ verified 2026-09-26
Important

Mixed Precision Training

Paulius Micikevicius, Sharan Narang et al. · 2017 · ICLR 2018

The recipe for training in 16-bit floating point without losing accuracy, which roughly halves activation memory and unlocks GPUs' fastest arithmetic.

Problem
FP16's narrow range makes small gradients round to zero and weight updates vanish.
What was new
Keep an FP32 master copy of the weights, scale the loss up so small gradients survive in FP16, and accumulate products in FP32.
~20 min readarXiv:1710.03740✓ verified 2026-10-04
Optional

A Study of BFLOAT16 for Deep Learning Training

Dhiraj Kalamkar, Dheevatsa Mudigere et al. · 2019

Showed that the bfloat16 format, with FP32's range and fewer digits of precision, trains networks to the same accuracy without loss scaling.

Problem
FP16 training needs loss scaling and careful tuning because its range is so narrow.
What was new
A broad study across vision, speech, language and recommendation models showing BF16 matches FP32 accuracy with no hyperparameter changes.
~15 min readarXiv:1905.12322✓ verified 2026-10-04
Essential

ZeRO: Memory Optimizations Toward Training Trillion Parameter Models

Samyam Rajbhandari, Jeff Rasley et al. · 2019

Explained where training memory goes (16 bytes per parameter with mixed-precision Adam) and how to remove the redundant copies. Its stages became DeepSpeed's ZeRO and PyTorch's FSDP.

Problem
Data parallelism keeps a full copy of the weights, gradients and optimizer states on every GPU, so models that don't fit on one GPU can't use it.
What was new
Shard the optimizer states, then the gradients, then the parameters across data-parallel GPUs, gathering each piece only when needed: memory per GPU falls in proportion to the number of GPUs.

How to read it: Section 3 ('Where did all the memory go?') and Figure 1 are the essentials; the rest is engineering detail.

~40 min readarXiv:1910.02054✓ verified 2026-10-04
Optional

Training Deep Nets with Sublinear Memory Cost

Tianqi Chen, Bing Xu et al. · 2016

Activation checkpointing: trade a little extra computation for a large cut in training memory. Every large model run uses some form of it.

Problem
Storing every intermediate activation for the backward pass makes memory grow linearly with depth.
What was new
Store activations only at checkpoints and recompute the rest during the backward pass: O(√n) memory for an n-layer network at the cost of about one extra forward pass.
~25 min readarXiv:1604.06174✓ verified 2026-10-04
Optional

Reducing Activation Recomputation in Large Transformer Models

Vijay Korthikanti, Jared Casper et al. · 2022

Worked out exactly how much activation memory a Transformer layer needs, and how to recompute only the cheap, memory-hungry parts.

Problem
Recomputing every layer saves memory but costs 30–40% more compute.
What was new
A per-layer activation formula, sequence parallelism, and selective recomputation of the attention scores, which cuts activation memory about fivefold with little overhead.

How to read it: Section 4.1 derives the 34 + 5as/h formula used in this chapter's memory calculator.

~30 min readarXiv:2205.05198✓ verified 2026-10-04
Important

Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Mohammad Shoeybi, Mostofa Patwary et al. · 2019

Introduced the tensor-parallel layout for Transformers that most large training systems still use: split each layer's matrices across GPUs.

Problem
Multi-billion-parameter Transformers no longer fit on a single GPU.
What was new
Split the attention heads and the feed-forward matrices across GPUs so each layer needs only a couple of all-reduces; trained an 8.3B model on 512 GPUs at 76% scaling efficiency.

How to read it: Figure 3 (how an MLP and an attention block are split) is the part to understand.

~25 min readarXiv:1909.08053✓ verified 2026-10-04
Optional

GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism

Yanping Huang, Youlong Cheng et al. · 2018

Made pipeline parallelism practical: split a network's layers across accelerators and keep them busy by streaming micro-batches through.

Problem
A model too deep for one accelerator can be split by layers, but then only one accelerator works at a time.
What was new
Split each batch into micro-batches and pipeline them through the stages; the idle 'bubble' becomes negligible once there are at least four times as many micro-batches as stages.
~25 min readarXiv:1811.06965✓ verified 2026-10-04
Important

Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Deepak Narayanan, Mohammad Shoeybi et al. · 2021

Showed how to compose tensor, pipeline and data parallelism to train trillion-parameter models efficiently on thousands of GPUs.

Problem
Each form of parallelism stops scaling on its own: tensor parallelism across servers is too slow, and pipelines waste time in bubbles.
What was new
Practical rules for combining the three (tensor parallelism inside a server, pipelines across servers) and an interleaved schedule: a 1-trillion-parameter model at 502 petaFLOP/s on 3,072 GPUs, 52% of peak.

How to read it: Read the takeaways in Section 3; they summarise how to choose the parallel sizes.

~45 min readarXiv:2104.04473✓ verified 2026-10-04
Optional

PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel

Yanli Zhao, Andrew Gu et al. · 2023 · PVLDB 16(12)

Describes PyTorch's built-in version of ZeRO-style sharding, the default way to train models that don't fit on one GPU.

Problem
ZeRO's ideas needed a general, efficient implementation inside the main deep-learning framework.
What was new
Fully sharded data parallelism in PyTorch, with the engineering for overlapping communication and computation and for mixed precision.
~30 min readarXiv:2304.11277✓ verified 2026-10-04
Important

PaLM: Scaling Language Modeling with Pathways

Aakanksha Chowdhery, Sharan Narang et al. · 2022

A 540-billion-parameter model whose report defined model FLOPs utilization (MFU) and described, unusually frankly, the loss spikes of a very large run.

Problem
How efficiently can a dense model be trained across thousands of accelerators, and how do you keep a giant run stable?
What was new
Training across two TPU v4 pods at 46.2% MFU; loss spikes (about 20 in the largest run) handled by restarting from an earlier checkpoint and skipping the surrounding batches.

How to read it: Sections 4 (training infrastructure) and 5.1 (training instability) are the parts for this chapter.

~1 h readarXiv:2204.02311✓ verified 2026-10-04
Important

OPT: Open Pre-trained Transformer Language Models

Susan Zhang, Stephen Roller et al. · 2022

Released a GPT-3-sized model with its full training logbook: a rare, honest record of what goes wrong in a large run.

Problem
How large models are actually trained, including failures and mid-run fixes, was rarely documented.
What was new
OPT-175B on 992 A100 GPUs, with at least 35 manual restarts from hardware failures over two months, and loss divergences handled by lowering the learning rate and restarting from earlier checkpoints.

How to read it: Section 2.5, 'Training Processes', and the released logbook.

~30 min readarXiv:2205.01068✓ verified 2026-10-04
Essential

Scaling Laws for Neural Language Models

Jared Kaplan, Sam McCandlish et al. · 2020

Found that language-model loss falls as a smooth power law in parameters, data and compute — making model scale something you could plan.

Problem
There was no quantitative way to predict how much better a larger model would be.
What was new
Empirical power-law fits of loss against model size, dataset size and compute, over many orders of magnitude.
~1 h readarXiv:2001.08361✓ verified 2026-09-26
Essential

Training Compute-Optimal Large Language Models

Jordan Hoffmann, Sebastian Borgeaud et al. · 2022 · NeurIPS 2022

Showed that many large models were undertrained: for a fixed compute budget, parameters and training tokens should grow roughly in proportion.

Problem
Earlier scaling recommendations favoured very large models trained on comparatively little data.
What was new
Trained 400+ models to fit compute-optimal trade-offs; the 70B 'Chinchilla' model outperformed much larger models trained on fewer tokens.
~1 h readarXiv:2203.15556✓ verified 2026-09-26
Important

Chinchilla Scaling: A replication attempt

Tamay Besiroglu, Ege Erdil et al. · 2024

Re-fitted Chinchilla's scaling law from its published data and found the printed coefficients inconsistent with the paper's own conclusions: a lesson in checking influential numbers.

Problem
Chinchilla's third method gave a fit that recommended far more tokens per parameter than its other two methods, with implausibly tight confidence intervals.
What was new
A re-fit (E = 1.82, A = 482, B = 2085, α = 0.348, β = 0.366) consistent with about 20 tokens per parameter, and an explanation: the original optimiser stopped early and the printed values were rounded.

How to read it: Short and readable; a good model of how to replicate a result from a figure.

~20 min readarXiv:2404.10102✓ verified 2026-10-04
Optional

Beyond Chinchilla-Optimal: Accounting for Inference in Language Model Scaling Laws

Nikhil Sardana, Jacob Portes et al. · 2023

Formalised why labs train models far past the 'compute-optimal' point: the model will be run billions of times.

Problem
Chinchilla's rule minimises training cost only, but a deployed model's inference cost can dwarf it.
What was new
Modified the Chinchilla laws to include inference demand; with large expected demand (around a billion requests), the cheapest model is smaller and trained on more data.
~25 min readarXiv:2401.00448✓ verified 2026-10-04
Essential

The Llama 3 Herd of Models

Aaron Grattafiori, Abhimanyu Dubey et al. · 2024

The most complete public account of building a frontier-scale model end to end: data pipeline, scaling-law experiments, 16,384-GPU training, failures and all.

Problem
Frontier labs had largely stopped publishing how their models were built.
What was new
A 405B model trained on 15.6T tokens with 3.8 × 10²⁵ FLOPs, with detailed data cleaning, a data mix chosen by scaling experiments, 4D parallelism at 38–43% utilisation, and 466 job interruptions in a 54-day period.

How to read it: It is 90+ pages. For this chapter read Section 3 (pre-training) only: data, scaling laws, infrastructure and the training recipe.

~2 h readarXiv:2407.21783✓ verified 2026-10-04

Watch

4 h 1 min

Andrej Karpathy

Let's reproduce GPT-2 (124M)

Build and train the smallest GPT-2 from scratch in PyTorch, then optimise it. For when you want to implement what this chapter describes.

Covers: The GPT-2 network in code, the engineering that makes training fast, and a full training run using the GPT-2 and GPT-3 papers' hyperparameters.

Frontier
3 h 31 min

Andrej Karpathy

Deep Dive into LLMs like ChatGPT

A long, general-audience walk through the whole pipeline behind a chat model, from internet text to tokens to pretraining to post-training.

Covers: The full training stack behind a chat model, from internet text and tokens to a pretrained base model and the post-training that turns it into an assistant.

Should know

What came next?

Chapter 10

From Base Model to Assistant →

A base model can answer questions, but pretraining alone does not make it reliably follow instructions or conversation roles.