Skip to content
Road to Intelligence

Concept · Chapter 9: How an LLM Is Actually Built

The Pretraining Loop

Must knowKnow well14 minDifficulty

Pretraining repeats one step about a million times: take a batch of millions of tokens, compute the average next-token loss, backpropagate, and let the optimizer nudge every weight, usually seeing each piece of data about once.

The problem

Chapter 4's training loop assumed a dataset you can pass over many times and a batch that fits on one device; neither holds when the data is trillions of tokens and each step processes millions.

The solution

Measure batches in tokens, accumulate gradients over micro-batches until the batch is complete, step the optimizer once per batch, mostly make a single pass over the data, and track loss on held-out text.

The consequence

Training becomes a long, mostly one-way stream: overfitting barely appears, the loss curve is the main health signal, and the batch size and learning rate become hyperparameters tied to the scale of the run.

One step, a million times

Chapter 4's loop still holds. A step is: run a batch forward, compute the average cross-entropy of predicting each next token, backpropagate, and let the optimizer (usually AdamW, a variant of Adam) update every weight. Pretraining just does it at a size where every word needs a number.

Batches are counted in tokens. GPT-3 175B used a batch of 3.2 million tokens Established. Llama 3 405B started with batches of 4 million tokens in sequences of 4,096, doubled to 8 million in sequences of 8,192, and doubled again to 16 million later in training Established.

Tiny example. A 16-million-token batch of 8,192-token sequences is 2,048 sequences. Training on 15.6 trillion tokens at that size would take 15.6 × 10¹² ÷ 1.6 × 10⁷ ≈ 975,000 steps. Llama 3 405B's learning-rate schedule spanned 1.2 million steps Established, in the same range because its early batches were smaller.

Gradient accumulation

No GPU holds 2,048 long sequences at once. So each GPU runs a few sequences (a micro-batch), keeps the gradients, runs a few more, and adds them up. Only when the whole batch has been seen does the optimizer step. The result is mathematically the same as one big batch; it just takes longer. Across many GPUs, data parallelism does the same thing in parallel.

Roughly one pass

With trillions of tokens, most data is seen once. Some valuable sources are repeated on purpose (see data mixtures), but there is no long sweep of epochs. Two consequences follow. Classic overfitting, where training loss falls while validation loss rises, rarely shows up; and the training loss on new batches is itself a fair estimate of generalisation, because the model has never seen those tokens.

Teams still track loss on held-out text from each source, and benchmark scores at checkpoints, to catch problems that the average hides.

Why should I care?

As a researcher

Batch size, learning rate and their schedules interact with model scale; getting them wrong wastes runs, and many 'architecture' results turn out to be tuning results.

As an engineer

Throughput is measured in tokens per second, cost in tokens processed. Knowing the step anatomy tells you where time and memory go.

Modern systems that depend on it

  • learning-rate schedules
  • gradient accumulation
  • data parallelism
  • the compute budget

Historical context

Before

Small datasets swept for many epochs, with batches of a few hundred examples and early stopping against overfitting.

After

Batches of millions of tokens, roughly one pass over trillions of tokens, and a loss curve watched for months.

Used today

Every pretraining run. Llama 3 405B's batch grew from 4M to 16M tokens over about 1.2 million optimizer steps.

What to remember

  • One step: forward pass on a batch → average loss → backward pass → optimizer update.
  • Batches are counted in tokens: GPT-3 175B used 3.2M; Llama 3 405B grew from 4M to 16M.
  • Gradient accumulation: sum gradients over several micro-batches, then take one step.
  • Pretraining usually makes about one pass over most data; repeated sources are a deliberate choice.
  • Watch the loss on held-out text: a smooth decline is healthy, a spike or plateau needs attention.

Key papers

Essential

Language Models are Few-Shot Learners

Tom B. Brown, Benjamin Mann et al. · 2020 · NeurIPS 2020

GPT-3 (175B parameters) showed that a large enough language model can perform new tasks from a few examples in its prompt, without any gradient updates.

How to read it: 75 pages. Sections 1–2 and Figure 1.2 carry the core idea; Section 6 on broader impacts is worth reading too.

~1 h 30 min readarXiv:2005.14165✓ verified 2026-09-26
Essential

Adam: A Method for Stochastic Optimization

Diederik P. Kingma, Jimmy Ba · 2014 · ICLR 2015

The default optimizer (with its AdamW variant) for training neural networks, including essentially all Transformers.

~40 min readarXiv:1412.6980✓ verified 2026-09-26
Important

Decoupled Weight Decay Regularization

Ilya Loshchilov, Frank Hutter · 2017 · ICLR 2019

Introduced AdamW, the variant of Adam used to train most modern Transformers and LLMs.

~40 min readarXiv:1711.05101✓ verified 2026-09-26
Essential

The Llama 3 Herd of Models

Aaron Grattafiori, Abhimanyu Dubey et al. · 2024

The most complete public account of building a frontier-scale model end to end: data pipeline, scaling-law experiments, 16,384-GPU training, failures and all.

How to read it: It is 90+ pages. For this chapter read Section 3 (pre-training) only: data, scaling laws, infrastructure and the training recipe.

~2 h readarXiv:2407.21783✓ verified 2026-10-04

Watch

4 h 1 min

Andrej Karpathy

Let's reproduce GPT-2 (124M)

Build and train the smallest GPT-2 from scratch in PyTorch, then optimise it. For when you want to implement what this chapter describes.

Frontier