Skip to content
Road to Intelligence

Concept · Chapter 9: How an LLM Is Actually Built

Mixed-Precision Training (FP16 and BF16)

Must knowUnderstand11 minDifficulty

Mixed-precision training does most arithmetic in 16-bit floating point, which is faster and halves activation memory, while keeping a 32-bit master copy of the weights so tiny updates are not lost.

The problem

32-bit arithmetic is slow and memory-hungry, but plain 16-bit numbers either can't represent very small gradients (FP16) or are too coarse to accumulate tiny weight updates.

The solution

Run the forward and backward passes in 16 bits, keep FP32 master weights and optimizer states for the update, and with FP16 scale the loss up so small gradients survive; BF16 keeps FP32's range and needs no scaling.

The consequence

Training runs on GPUs' fastest units at roughly half the activation memory; BF16 became the default for LLM pretraining, and the push to even lower precision (FP8) continues.

Three ways to store a number

A floating-point number spends some bits on the exponent (how big or small it can be) and some on the fraction (how many significant digits it keeps).

FormatBytesExponent bitsFraction bitsRangePrecision
FP324823about 10⁻³⁸ to 10³⁸about 7 decimal digits
FP162510up to 65,504about 3 digits
BF16287same as FP32about 2 digits

What goes wrong in plain 16 bits

Underflow. Gradients are often tiny. In FP16, anything much below about 6 × 10⁻⁸ becomes zero, and that part of the signal is lost.

Swamped updates. A weight of 1.0 and an update of 0.0001: in a format with only 2–3 significant digits, 1.0 + 0.0001 rounds back to 1.0. Thousands of small steps that should add up do nothing.

The recipe

Micikevicius and colleagues kept an FP32 master copy of the weights, scaled the loss up so small gradients stay representable in FP16, and accumulated products in FP32; with these three techniques, half-precision training matched FP32 accuracy across many models Established.

Tiny example. A gradient of 10⁻⁸ underflows in FP16. Multiply the loss by 1,024 before the backward pass and the gradient becomes about 10⁻⁵, which FP16 can hold. Divide by 1,024 again, in FP32, before the update.

BF16 sidesteps the range problem. BFLOAT16 has the same range as FP32, so models can be trained in it without loss scaling or hyperparameter changes Established. Its precision is lower than FP16's, which the FP32 master weights absorb. Llama 3's training efficiency was reported as BF16 model FLOPs utilisation Established, and BF16 is the usual choice for pretraining today.

What it saves, and what it doesn't

Activations, the intermediate values kept for the backward pass, are stored in 16 bits: half the memory. And GPUs run 16-bit matrix multiplication much faster than 32-bit.

But with Adam the model states barely shrink. Mixed-precision Adam keeps 16-bit weights and gradients (2 + 2 bytes per parameter) plus an FP32 copy of the weights and both Adam moments (4 + 4 + 4 bytes): 16 bytes per parameter Established, the same as plain FP32 training's 4 + 4 + 8. See where the memory goes.

What to remember

  • FP32: 8 exponent bits, 23 fraction bits. FP16: 5 and 10. BF16: 8 and 7.
  • FP16's largest value is 65,504; small gradients underflow to zero without loss scaling.
  • BF16 has FP32's range with less precision, so it usually needs no loss scaling.
  • Master weights stay in FP32 so that small updates add up.
  • With Adam, model states still cost about 16 bytes per parameter; the savings are in activations and speed.

Key papers

Important

Mixed Precision Training

Paulius Micikevicius, Sharan Narang et al. · 2017 · ICLR 2018

The recipe for training in 16-bit floating point without losing accuracy, which roughly halves activation memory and unlocks GPUs' fastest arithmetic.

~20 min readarXiv:1710.03740✓ verified 2026-10-04
Optional

A Study of BFLOAT16 for Deep Learning Training

Dhiraj Kalamkar, Dheevatsa Mudigere et al. · 2019

Showed that the bfloat16 format, with FP32's range and fewer digits of precision, trains networks to the same accuracy without loss scaling.

~15 min readarXiv:1905.12322✓ verified 2026-10-04
Essential

ZeRO: Memory Optimizations Toward Training Trillion Parameter Models

Samyam Rajbhandari, Jeff Rasley et al. · 2019

Explained where training memory goes (16 bytes per parameter with mixed-precision Adam) and how to remove the redundant copies. Its stages became DeepSpeed's ZeRO and PyTorch's FSDP.

How to read it: Section 3 ('Where did all the memory go?') and Figure 1 are the essentials; the rest is engineering detail.

~40 min readarXiv:1910.02054✓ verified 2026-10-04

Watch