Skip to content
Road to Intelligence

Concept · Chapter 9: How an LLM Is Actually Built

Where Training Memory Goes

Must knowKnow well15 minDifficulty

Training needs memory for the weights, their gradients, the optimizer's states (about 16 bytes per parameter in total with mixed-precision Adam) and the activations saved for the backward pass, which grow with sequence length and batch size.

The problem

A 7-billion-parameter model's 16-bit weights take 14 GB, which fits on one GPU, yet training it there is impossible. Where does the memory go?

The solution

Account for each part: shard the model states across GPUs, recompute activations instead of storing them, and shrink micro-batches, until each GPU's share fits.

The consequence

Memory, not arithmetic, decides how a model must be split across GPUs; the techniques that save it (sharding, recomputation, parallelism) each cost communication or extra compute.

The puzzle

The ZeRO paper opens with this puzzle: a 1.5-billion-parameter GPT-2 needs only 3 GB for its 16-bit weights, yet could not be trained on a single 32 GB GPU Established. The weights are the smallest part of the bill.

Model states: 16 bytes per parameter

With mixed-precision Adam, each parameter carries:

WhatPrecisionBytes
Weight used in forward and backward16-bit2
Gradient16-bit2
Master copy of the weight32-bit4
Adam's first moment (momentum)32-bit4
Adam's second moment (variance)32-bit4
Total16

That is ZeRO's accounting: 2 + 2 + 12 = 16 bytes per parameter, so the 1.5B GPT-2 needs at least 24 GB just for model states Established.

Tiny example. LLaMA 7B has 6.7 billion parameters. Weights for inference: 6.7 × 10⁹ × 2 bytes ≈ 13.4 GB. Model states for training: 6.7 × 10⁹ × 16 ≈ 107 GB. An 80 GB GPU can run it but cannot train it, before a single activation is stored.

Activations: the part that grows with the data

Backpropagation needs the intermediate values from the forward pass. Their size grows with the sequence length s, the micro-batch b, the hidden size h and the number of layers L. Korthikanti and colleagues derived the activation memory of one GPT-style Transformer layer in 16-bit as sbh(34 + 5as/h) bytes, where a is the number of attention heads Established. The second term is the attention scores, and it grows with the square of the sequence length.

For LLaMA 7B (h = 4,096, a = 32, 32 layers) with one 4,096-token sequence, that is about 3.3 GB per layer, or 104 GB in total: as much as the model states.

Trading compute for memory

Recomputation (activation checkpointing) keeps only some activations and recomputes the rest during the backward pass. Chen and colleagues showed that an n-layer network can be trained with O(√n) activation memory for the cost of about one extra forward pass Established. Storing only each layer's input cuts LLaMA 7B's example to about 1 GB, but full recomputation added 30–40% to training time in Korthikanti and colleagues' runs Established. Their alternative, recomputing only the attention scores, reduced activation memory about fivefold while removing over 90% of the recomputation overhead Established.

Try the combinations in the calculator: the model states shrink only by spreading them across GPUs (sharded training), and the activations shrink only by recomputing, shortening, or splitting the layers themselves (model parallelism).

Try it

Will It Fit?

Add up the memory a GPU needs to train a real model: weights, gradients, optimizer states and activations. Then shard and recompute until it fits in 80 GB.

Know well8 min

Why should I care?

As a researcher

Memory limits decide which experiments are feasible on the hardware you have: sequence length, batch size and model size trade against each other.

As an engineer

'CUDA out of memory' is the most common training failure. Knowing the budget lets you fix it deliberately (shard, recompute, shorten) instead of by trial and error.

Modern systems that depend on it

  • ZeRO and FSDP
  • tensor and pipeline parallelism
  • context length
  • fine-tuning on one GPU

Historical context

Before

Models were small enough that the weights, gradients and activations all fitted on one GPU, and memory was rarely the question.

After

Training plans start from a memory budget: 16 bytes per parameter of model states, plus activations that are recomputed, sharded or split across GPUs.

Used today

Every large training run combines sharded model states with activation recomputation; fine-tuning on a single GPU depends on the same accounting (and on tricks from Chapter 11).

What to remember

  • Mixed-precision Adam: 2 (weights) + 2 (gradients) + 12 (FP32 weights and two moments) = 16 bytes per parameter.
  • So LLaMA 7B needs about 107 GB of model states before any activations: more than one 80 GB GPU.
  • Activations grow with sequence length × batch × layers; storing every one can exceed the model states.
  • Recomputation stores only checkpoints and recomputes the rest in the backward pass.
  • Full recomputation costs about one extra forward pass (30–40% more time); selective recomputation much less.

Key papers

Optional

Training Deep Nets with Sublinear Memory Cost

Tianqi Chen, Bing Xu et al. · 2016

Activation checkpointing: trade a little extra computation for a large cut in training memory. Every large model run uses some form of it.

~25 min readarXiv:1604.06174✓ verified 2026-10-04
Important

Mixed Precision Training

Paulius Micikevicius, Sharan Narang et al. · 2017 · ICLR 2018

The recipe for training in 16-bit floating point without losing accuracy, which roughly halves activation memory and unlocks GPUs' fastest arithmetic.

~20 min readarXiv:1710.03740✓ verified 2026-10-04
Essential

ZeRO: Memory Optimizations Toward Training Trillion Parameter Models

Samyam Rajbhandari, Jeff Rasley et al. · 2019

Explained where training memory goes (16 bytes per parameter with mixed-precision Adam) and how to remove the redundant copies. Its stages became DeepSpeed's ZeRO and PyTorch's FSDP.

How to read it: Section 3 ('Where did all the memory go?') and Figure 1 are the essentials; the rest is engineering detail.

~40 min readarXiv:1910.02054✓ verified 2026-10-04
Optional

Reducing Activation Recomputation in Large Transformer Models

Vijay Korthikanti, Jared Casper et al. · 2022

Worked out exactly how much activation memory a Transformer layer needs, and how to recompute only the cheap, memory-hungry parts.

How to read it: Section 4.1 derives the 34 + 5as/h formula used in this chapter's memory calculator.

~30 min readarXiv:2205.05198✓ verified 2026-10-04

Watch