Concept · Chapter 9: How an LLM Is Actually Built
Loss Spikes, Failures and Checkpoints
Long training runs break: the loss can suddenly spike or diverge, and among thousands of GPUs something fails every few hours, so runs save checkpoints often and are built to roll back and resume.
The problem
A synchronous job on thousands of GPUs runs for months; a single failing GPU can stop it, and an unstable optimisation can wreck weeks of progress in minutes.
The solution
Save checkpoints frequently, detect failures and restart automatically, and handle divergence by rolling back to an earlier checkpoint, skipping the offending batches or lowering the learning rate.
The consequence
Reliability engineering became part of model training: the fraction of time spent actually training ('effective training time') is a key metric of a large run.
You should understand first
- Derivatives and Gradients
- Loss Functions
- Gradient Descent
- Probability and Distributions
- Expected Value and Variance
- Stochastic Gradient Descent (SGD)
- The Chain Rule
- Vectors
- Dot Product
- The Turing Test
- Symbolic AI
- Logic and Rules
- Expert Systems
- Knowledge Representation
- The Knowledge-Acquisition Bottleneck
- From Rules to Learning
- Supervised, Unsupervised and Self-Supervised Learning
- Features, Labels and Tasks
- Linear Regression
- Entropy
- Softmax
- Cross-Entropy Loss
- Logistic Regression
- The Perceptron
- Activation Functions
- The Artificial Neuron
- Matrix Multiplication
- Multilayer Perceptron (MLP)
- The Forward Pass
- Computational Graphs and Autodiff
- Backpropagation
- Momentum and Adam
- The Pretraining Loop
- Learning-Rate Schedules, Warmup and Clipping
- Loss Spikes, Failures and Checkpoints
Loss spikes
Sometimes the loss, falling smoothly for weeks, jumps up. PaLM's team saw about 20 such spikes while training their largest model, despite gradient clipping, at irregular times. Their fix was to restart from a checkpoint about 100 steps before the spike and skip roughly 200–500 batches of data Established. They did not think the spikes were caused by bad data alone: the same batches, replayed from an earlier checkpoint, did not cause a spike Established. What causes loss spikes in very large models is still not fully understood Active research.
OPT-175B's team handled loss divergences by lowering the learning rate and restarting from an earlier checkpoint Established, and published the logbook of every intervention.
Hardware fails, constantly
At this scale failure is routine. Training OPT-175B on 992 GPUs involved at least 35 manual restarts and cycling out over 100 faulty machines in two months Established. During a 54-day period of Llama 3 405B pretraining there were 466 job interruptions: 47 planned and 419 unexpected, about 78% of them attributed to confirmed or suspected hardware problems, most often GPUs Established. Even so, automation kept effective training time above 90%; significant manual intervention was needed only three times Established.
Tiny example. 419 unexpected interruptions in 54 days is about one every 3 hours. If restarting from the last checkpoint loses on average 15 minutes of progress plus 10 minutes of restart time, that is roughly 175 hours lost: about 13% of the period. Hence the obsession with fast checkpoints and fast restarts. (The 15 and 10 minutes are illustrative.)
Checkpoints
A checkpoint saves the weights, optimizer states and data position so training can resume exactly. For a 405B model the optimizer states alone are terabytes, so writing them quickly is a storage-engineering problem in its own right.
What to remember
- A loss spike is a sudden jump in the training loss; sometimes it recovers, sometimes the run diverges.
- PaLM 540B: about 20 spikes; fix = restart ~100 steps earlier and skip 200–500 batches.
- OPT-175B: at least 35 manual restarts from hardware failures in two months.
- Llama 3 405B: 466 interruptions in 54 days, about 78% of the unexpected ones hardware-related, yet over 90% effective training time.
Key papers
PaLM: Scaling Language Modeling with Pathways
Aakanksha Chowdhery, Sharan Narang et al. · 2022
A 540-billion-parameter model whose report defined model FLOPs utilization (MFU) and described, unusually frankly, the loss spikes of a very large run.
How to read it: Sections 4 (training infrastructure) and 5.1 (training instability) are the parts for this chapter.
OPT: Open Pre-trained Transformer Language Models
Susan Zhang, Stephen Roller et al. · 2022
Released a GPT-3-sized model with its full training logbook: a rare, honest record of what goes wrong in a large run.
How to read it: Section 2.5, 'Training Processes', and the released logbook.
The Llama 3 Herd of Models
Aaron Grattafiori, Abhimanyu Dubey et al. · 2024
The most complete public account of building a frontier-scale model end to end: data pipeline, scaling-law experiments, 16,384-GPU training, failures and all.
How to read it: It is 90+ pages. For this chapter read Section 3 (pre-training) only: data, scaling laws, infrastructure and the training recipe.