Concept · Chapter 9: How an LLM Is Actually Built
The Pretraining Loop
Pretraining repeats one step about a million times: take a batch of millions of tokens, compute the average next-token loss, backpropagate, and let the optimizer nudge every weight, usually seeing each piece of data about once.
The problem
Chapter 4's training loop assumed a dataset you can pass over many times and a batch that fits on one device; neither holds when the data is trillions of tokens and each step processes millions.
The solution
Measure batches in tokens, accumulate gradients over micro-batches until the batch is complete, step the optimizer once per batch, mostly make a single pass over the data, and track loss on held-out text.
The consequence
Training becomes a long, mostly one-way stream: overfitting barely appears, the loss curve is the main health signal, and the batch size and learning rate become hyperparameters tied to the scale of the run.
You should understand first
- Derivatives and Gradients
- Loss Functions
- Gradient Descent
- Probability and Distributions
- Expected Value and Variance
- Stochastic Gradient Descent (SGD)
- The Chain Rule
- Vectors
- Dot Product
- The Turing Test
- Symbolic AI
- Logic and Rules
- Expert Systems
- Knowledge Representation
- The Knowledge-Acquisition Bottleneck
- From Rules to Learning
- Supervised, Unsupervised and Self-Supervised Learning
- Features, Labels and Tasks
- Linear Regression
- Entropy
- Softmax
- Cross-Entropy Loss
- Logistic Regression
- The Perceptron
- Activation Functions
- The Artificial Neuron
- Matrix Multiplication
- Multilayer Perceptron (MLP)
- The Forward Pass
- Computational Graphs and Autodiff
- Backpropagation
- Momentum and Adam
- The Pretraining Loop
One step, a million times
Chapter 4's loop still holds. A step is: run a batch forward, compute the average cross-entropy of predicting each next token, backpropagate, and let the optimizer (usually AdamW, a variant of Adam) update every weight. Pretraining just does it at a size where every word needs a number.
Batches are counted in tokens. GPT-3 175B used a batch of 3.2 million tokens Established. Llama 3 405B started with batches of 4 million tokens in sequences of 4,096, doubled to 8 million in sequences of 8,192, and doubled again to 16 million later in training Established.
Tiny example. A 16-million-token batch of 8,192-token sequences is 2,048 sequences. Training on 15.6 trillion tokens at that size would take 15.6 × 10¹² ÷ 1.6 × 10⁷ ≈ 975,000 steps. Llama 3 405B's learning-rate schedule spanned 1.2 million steps Established, in the same range because its early batches were smaller.
Gradient accumulation
No GPU holds 2,048 long sequences at once. So each GPU runs a few sequences (a micro-batch), keeps the gradients, runs a few more, and adds them up. Only when the whole batch has been seen does the optimizer step. The result is mathematically the same as one big batch; it just takes longer. Across many GPUs, data parallelism does the same thing in parallel.
Roughly one pass
With trillions of tokens, most data is seen once. Some valuable sources are repeated on purpose (see data mixtures), but there is no long sweep of epochs. Two consequences follow. Classic overfitting, where training loss falls while validation loss rises, rarely shows up; and the training loss on new batches is itself a fair estimate of generalisation, because the model has never seen those tokens.
Teams still track loss on held-out text from each source, and benchmark scores at checkpoints, to catch problems that the average hides.
Why should I care?
As a researcher
Batch size, learning rate and their schedules interact with model scale; getting them wrong wastes runs, and many 'architecture' results turn out to be tuning results.
As an engineer
Throughput is measured in tokens per second, cost in tokens processed. Knowing the step anatomy tells you where time and memory go.
Modern systems that depend on it
- learning-rate schedules
- gradient accumulation
- data parallelism
- the compute budget
Historical context
Before
Small datasets swept for many epochs, with batches of a few hundred examples and early stopping against overfitting.
After
Batches of millions of tokens, roughly one pass over trillions of tokens, and a loss curve watched for months.
Used today
Every pretraining run. Llama 3 405B's batch grew from 4M to 16M tokens over about 1.2 million optimizer steps.
What to remember
- One step: forward pass on a batch → average loss → backward pass → optimizer update.
- Batches are counted in tokens: GPT-3 175B used 3.2M; Llama 3 405B grew from 4M to 16M.
- Gradient accumulation: sum gradients over several micro-batches, then take one step.
- Pretraining usually makes about one pass over most data; repeated sources are a deliberate choice.
- Watch the loss on held-out text: a smooth decline is healthy, a spike or plateau needs attention.
Key papers
Language Models are Few-Shot Learners
Tom B. Brown, Benjamin Mann et al. · 2020 · NeurIPS 2020
GPT-3 (175B parameters) showed that a large enough language model can perform new tasks from a few examples in its prompt, without any gradient updates.
How to read it: 75 pages. Sections 1–2 and Figure 1.2 carry the core idea; Section 6 on broader impacts is worth reading too.
Adam: A Method for Stochastic Optimization
Diederik P. Kingma, Jimmy Ba · 2014 · ICLR 2015
The default optimizer (with its AdamW variant) for training neural networks, including essentially all Transformers.
Decoupled Weight Decay Regularization
Ilya Loshchilov, Frank Hutter · 2017 · ICLR 2019
Introduced AdamW, the variant of Adam used to train most modern Transformers and LLMs.
The Llama 3 Herd of Models
Aaron Grattafiori, Abhimanyu Dubey et al. · 2024
The most complete public account of building a frontier-scale model end to end: data pipeline, scaling-law experiments, 16,384-GPU training, failures and all.
How to read it: It is 90+ pages. For this chapter read Section 3 (pre-training) only: data, scaling laws, infrastructure and the training recipe.
Watch
Andrej Karpathy
Let's reproduce GPT-2 (124M)
Build and train the smallest GPT-2 from scratch in PyTorch, then optimise it. For when you want to implement what this chapter describes.
Stanford Online
Stanford CS336 Language Modeling from Scratch | Spring 2025 | Lec. 2: Pytorch, Resource Accounting
Teaches the habit this chapter is built on: napkin maths for memory (bytes per parameter) and compute (FLOPs) before you train anything.