Concept · Chapter 9: How an LLM Is Actually Built
Where Training Memory Goes
Training needs memory for the weights, their gradients, the optimizer's states (about 16 bytes per parameter in total with mixed-precision Adam) and the activations saved for the backward pass, which grow with sequence length and batch size.
The problem
A 7-billion-parameter model's 16-bit weights take 14 GB, which fits on one GPU, yet training it there is impossible. Where does the memory go?
The solution
Account for each part: shard the model states across GPUs, recompute activations instead of storing them, and shrink micro-batches, until each GPU's share fits.
The consequence
Memory, not arithmetic, decides how a model must be split across GPUs; the techniques that save it (sharding, recomputation, parallelism) each cost communication or extra compute.
You should understand first
- Derivatives and Gradients
- Loss Functions
- Gradient Descent
- Probability and Distributions
- Expected Value and Variance
- Stochastic Gradient Descent (SGD)
- The Chain Rule
- Vectors
- Dot Product
- The Turing Test
- Symbolic AI
- Logic and Rules
- Expert Systems
- Knowledge Representation
- The Knowledge-Acquisition Bottleneck
- From Rules to Learning
- Supervised, Unsupervised and Self-Supervised Learning
- Features, Labels and Tasks
- Linear Regression
- Entropy
- Softmax
- Cross-Entropy Loss
- Logistic Regression
- The Perceptron
- Activation Functions
- The Artificial Neuron
- Matrix Multiplication
- Multilayer Perceptron (MLP)
- The Forward Pass
- Computational Graphs and Autodiff
- Backpropagation
- Momentum and Adam
- The Pretraining Loop
- Mixed-Precision Training (FP16 and BF16)
- Text as Data
- One-Hot Encoding
- Tokenization
- Conditional Probability and Bayes' Theorem
- Probability of Sequences
- Language Modeling
- Embeddings
- Attention
- Self-Attention
- Causal Masking
- Autoregressive Next-Token Prediction
- Pretraining at Scale
- Parameters, Tokens and Context Windows
- Where Training Memory Goes
The puzzle
The ZeRO paper opens with this puzzle: a 1.5-billion-parameter GPT-2 needs only 3 GB for its 16-bit weights, yet could not be trained on a single 32 GB GPU Established. The weights are the smallest part of the bill.
Model states: 16 bytes per parameter
With mixed-precision Adam, each parameter carries:
| What | Precision | Bytes |
|---|---|---|
| Weight used in forward and backward | 16-bit | 2 |
| Gradient | 16-bit | 2 |
| Master copy of the weight | 32-bit | 4 |
| Adam's first moment (momentum) | 32-bit | 4 |
| Adam's second moment (variance) | 32-bit | 4 |
| Total | 16 |
That is ZeRO's accounting: 2 + 2 + 12 = 16 bytes per parameter, so the 1.5B GPT-2 needs at least 24 GB just for model states Established.
Tiny example. LLaMA 7B has 6.7 billion parameters. Weights for inference: 6.7 × 10⁹ × 2 bytes ≈ 13.4 GB. Model states for training: 6.7 × 10⁹ × 16 ≈ 107 GB. An 80 GB GPU can run it but cannot train it, before a single activation is stored.
Activations: the part that grows with the data
Backpropagation needs the intermediate values from the forward pass. Their size grows with the sequence length s, the micro-batch b, the hidden size h and the number of layers L. Korthikanti and colleagues derived the activation memory of one GPT-style Transformer layer in 16-bit as sbh(34 + 5as/h) bytes, where a is the number of attention heads Established. The second term is the attention scores, and it grows with the square of the sequence length.
For LLaMA 7B (h = 4,096, a = 32, 32 layers) with one 4,096-token sequence, that is about 3.3 GB per layer, or 104 GB in total: as much as the model states.
Trading compute for memory
Recomputation (activation checkpointing) keeps only some activations and recomputes the rest during the backward pass. Chen and colleagues showed that an n-layer network can be trained with O(√n) activation memory for the cost of about one extra forward pass Established. Storing only each layer's input cuts LLaMA 7B's example to about 1 GB, but full recomputation added 30–40% to training time in Korthikanti and colleagues' runs Established. Their alternative, recomputing only the attention scores, reduced activation memory about fivefold while removing over 90% of the recomputation overhead Established.
Try the combinations in the calculator: the model states shrink only by spreading them across GPUs (sharded training), and the activations shrink only by recomputing, shortening, or splitting the layers themselves (model parallelism).
Try it
Add up the memory a GPU needs to train a real model: weights, gradients, optimizer states and activations. Then shard and recompute until it fits in 80 GB.
Why should I care?
As a researcher
Memory limits decide which experiments are feasible on the hardware you have: sequence length, batch size and model size trade against each other.
As an engineer
'CUDA out of memory' is the most common training failure. Knowing the budget lets you fix it deliberately (shard, recompute, shorten) instead of by trial and error.
Modern systems that depend on it
- ZeRO and FSDP
- tensor and pipeline parallelism
- context length
- fine-tuning on one GPU
Historical context
Before
Models were small enough that the weights, gradients and activations all fitted on one GPU, and memory was rarely the question.
After
Training plans start from a memory budget: 16 bytes per parameter of model states, plus activations that are recomputed, sharded or split across GPUs.
Used today
Every large training run combines sharded model states with activation recomputation; fine-tuning on a single GPU depends on the same accounting (and on tricks from Chapter 11).
What to remember
- Mixed-precision Adam: 2 (weights) + 2 (gradients) + 12 (FP32 weights and two moments) = 16 bytes per parameter.
- So LLaMA 7B needs about 107 GB of model states before any activations: more than one 80 GB GPU.
- Activations grow with sequence length × batch × layers; storing every one can exceed the model states.
- Recomputation stores only checkpoints and recomputes the rest in the backward pass.
- Full recomputation costs about one extra forward pass (30–40% more time); selective recomputation much less.
Key papers
Training Deep Nets with Sublinear Memory Cost
Tianqi Chen, Bing Xu et al. · 2016
Activation checkpointing: trade a little extra computation for a large cut in training memory. Every large model run uses some form of it.
Mixed Precision Training
Paulius Micikevicius, Sharan Narang et al. · 2017 · ICLR 2018
The recipe for training in 16-bit floating point without losing accuracy, which roughly halves activation memory and unlocks GPUs' fastest arithmetic.
ZeRO: Memory Optimizations Toward Training Trillion Parameter Models
Samyam Rajbhandari, Jeff Rasley et al. · 2019
Explained where training memory goes (16 bytes per parameter with mixed-precision Adam) and how to remove the redundant copies. Its stages became DeepSpeed's ZeRO and PyTorch's FSDP.
How to read it: Section 3 ('Where did all the memory go?') and Figure 1 are the essentials; the rest is engineering detail.
Reducing Activation Recomputation in Large Transformer Models
Vijay Korthikanti, Jared Casper et al. · 2022
Worked out exactly how much activation memory a Transformer layer needs, and how to recompute only the cheap, memory-hungry parts.
How to read it: Section 4.1 derives the 34 + 5as/h formula used in this chapter's memory calculator.
Watch
Stanford Online
Stanford CS336 Language Modeling from Scratch | Spring 2025 | Lec. 2: Pytorch, Resource Accounting
Teaches the habit this chapter is built on: napkin maths for memory (bytes per parameter) and compute (FLOPs) before you train anything.
Stanford Online
Stanford CS336 Language Modeling from Scratch | Spring 2025 | Lecture 7: Parallelism 1
A clear university lecture on how training is split across many GPUs, with the trade-offs between the methods.