Concept · Chapter 9: How an LLM Is Actually Built
Mixed-Precision Training (FP16 and BF16)
Mixed-precision training does most arithmetic in 16-bit floating point, which is faster and halves activation memory, while keeping a 32-bit master copy of the weights so tiny updates are not lost.
The problem
32-bit arithmetic is slow and memory-hungry, but plain 16-bit numbers either can't represent very small gradients (FP16) or are too coarse to accumulate tiny weight updates.
The solution
Run the forward and backward passes in 16 bits, keep FP32 master weights and optimizer states for the update, and with FP16 scale the loss up so small gradients survive; BF16 keeps FP32's range and needs no scaling.
The consequence
Training runs on GPUs' fastest units at roughly half the activation memory; BF16 became the default for LLM pretraining, and the push to even lower precision (FP8) continues.
You should understand first
- Derivatives and Gradients
- Loss Functions
- Gradient Descent
- Probability and Distributions
- Expected Value and Variance
- Stochastic Gradient Descent (SGD)
- The Chain Rule
- Vectors
- Dot Product
- The Turing Test
- Symbolic AI
- Logic and Rules
- Expert Systems
- Knowledge Representation
- The Knowledge-Acquisition Bottleneck
- From Rules to Learning
- Supervised, Unsupervised and Self-Supervised Learning
- Features, Labels and Tasks
- Linear Regression
- Entropy
- Softmax
- Cross-Entropy Loss
- Logistic Regression
- The Perceptron
- Activation Functions
- The Artificial Neuron
- Matrix Multiplication
- Multilayer Perceptron (MLP)
- The Forward Pass
- Computational Graphs and Autodiff
- Backpropagation
- Momentum and Adam
- The Pretraining Loop
- Mixed-Precision Training (FP16 and BF16)
Three ways to store a number
A floating-point number spends some bits on the exponent (how big or small it can be) and some on the fraction (how many significant digits it keeps).
| Format | Bytes | Exponent bits | Fraction bits | Range | Precision |
|---|---|---|---|---|---|
| FP32 | 4 | 8 | 23 | about 10⁻³⁸ to 10³⁸ | about 7 decimal digits |
| FP16 | 2 | 5 | 10 | up to 65,504 | about 3 digits |
| BF16 | 2 | 8 | 7 | same as FP32 | about 2 digits |
What goes wrong in plain 16 bits
Underflow. Gradients are often tiny. In FP16, anything much below about 6 × 10⁻⁸ becomes zero, and that part of the signal is lost.
Swamped updates. A weight of 1.0 and an update of 0.0001: in a format with only 2–3 significant digits, 1.0 + 0.0001 rounds back to 1.0. Thousands of small steps that should add up do nothing.
The recipe
Micikevicius and colleagues kept an FP32 master copy of the weights, scaled the loss up so small gradients stay representable in FP16, and accumulated products in FP32; with these three techniques, half-precision training matched FP32 accuracy across many models Established.
Tiny example. A gradient of 10⁻⁸ underflows in FP16. Multiply the loss by 1,024 before the backward pass and the gradient becomes about 10⁻⁵, which FP16 can hold. Divide by 1,024 again, in FP32, before the update.
BF16 sidesteps the range problem. BFLOAT16 has the same range as FP32, so models can be trained in it without loss scaling or hyperparameter changes Established. Its precision is lower than FP16's, which the FP32 master weights absorb. Llama 3's training efficiency was reported as BF16 model FLOPs utilisation Established, and BF16 is the usual choice for pretraining today.
What it saves, and what it doesn't
Activations, the intermediate values kept for the backward pass, are stored in 16 bits: half the memory. And GPUs run 16-bit matrix multiplication much faster than 32-bit.
But with Adam the model states barely shrink. Mixed-precision Adam keeps 16-bit weights and gradients (2 + 2 bytes per parameter) plus an FP32 copy of the weights and both Adam moments (4 + 4 + 4 bytes): 16 bytes per parameter Established, the same as plain FP32 training's 4 + 4 + 8. See where the memory goes.
What to remember
- FP32: 8 exponent bits, 23 fraction bits. FP16: 5 and 10. BF16: 8 and 7.
- FP16's largest value is 65,504; small gradients underflow to zero without loss scaling.
- BF16 has FP32's range with less precision, so it usually needs no loss scaling.
- Master weights stay in FP32 so that small updates add up.
- With Adam, model states still cost about 16 bytes per parameter; the savings are in activations and speed.
Key papers
Mixed Precision Training
Paulius Micikevicius, Sharan Narang et al. · 2017 · ICLR 2018
The recipe for training in 16-bit floating point without losing accuracy, which roughly halves activation memory and unlocks GPUs' fastest arithmetic.
A Study of BFLOAT16 for Deep Learning Training
Dhiraj Kalamkar, Dheevatsa Mudigere et al. · 2019
Showed that the bfloat16 format, with FP32's range and fewer digits of precision, trains networks to the same accuracy without loss scaling.
ZeRO: Memory Optimizations Toward Training Trillion Parameter Models
Samyam Rajbhandari, Jeff Rasley et al. · 2019
Explained where training memory goes (16 bytes per parameter with mixed-precision Adam) and how to remove the redundant copies. Its stages became DeepSpeed's ZeRO and PyTorch's FSDP.
How to read it: Section 3 ('Where did all the memory go?') and Figure 1 are the essentials; the rest is engineering detail.
Watch
Stanford Online
Stanford CS336 Language Modeling from Scratch | Spring 2025 | Lec. 2: Pytorch, Resource Accounting
Teaches the habit this chapter is built on: napkin maths for memory (bytes per parameter) and compute (FLOPs) before you train anything.