Skip to content
Road to Intelligence

Concept · Chapter 9: How an LLM Is Actually Built

Learning-Rate Schedules, Warmup and Clipping

Must knowKnow well10 minDifficulty

The learning rate is ramped up from near zero during a short warmup, then decayed (usually along a cosine curve) to a small fraction of its peak, while gradient clipping caps any single step that would be too large.

The problem

At the start of training the weights are random and Adam's statistics are unreliable, so large steps can blow the run up; near the end, large steps keep the loss bouncing instead of settling.

The solution

Warm up linearly over the first few thousand steps, decay smoothly afterwards, use AdamW's decoupled weight decay, and clip the gradient norm so one bad batch can't throw the weights far.

The consequence

A stable, nearly standard recipe that works across scales, with a peak learning rate that shrinks as models grow.

Warm up, then cool down

The learning rate (Chapter 2's step size) is not held fixed. A typical schedule has two parts:

  • Warmup. Start near zero and rise linearly to the peak over the first few thousand steps. Early on, the weights are random and Adam's running averages of the gradient have barely started, so a full-size step can push the model somewhere it never recovers from. Goyal and colleagues used a gradual warmup to train with very large batches without losing accuracy Established.
  • Decay. After the peak, lower the rate smoothly so the model can settle into a good region. The cosine schedule follows half a cosine curve from the peak down to a minimum Established:
ηt=ηmin⁡+12(ηmax⁡−ηmin⁡)(1+cos⁡πtT)\eta_t=\eta_{\min}+\tfrac12(\eta_{\max}-\eta_{\min})\left(1+\cos\frac{\pi t}{T}\right)

Tiny example. Halfway through the decay, cos⁡(π/2)=0\cos(\pi/2)=0, so the rate is exactly midway between peak and floor. With Llama 3 405B's numbers that is 8×10−7+12(8×10−5−8×10−7)≈4.04×10−58\times10^{-7}+\tfrac12(8\times10^{-5}-8\times10^{-7})\approx 4.04\times10^{-5}.

Learning rate over one real run · Llama 3 405B
0300k600k900k1200k04e-58e-5the whole run (steps)08k20k40k04e-58e-5the first 40,000 steps

A linear warmup to 8 × 10⁻⁵ over the first 8,000 steps (dashed line), then a cosine decay to 8 × 10⁻⁷ by step 1.2 million. The warmup is under 1% of the run, too short to see on the left.

Llama 3's peak learning rates fell as models grew: 3 × 10⁻⁴ for 8B, 1.5 × 10⁻⁴ for 70B and 8 × 10⁻⁵ for 405B Established. Larger models are more sensitive to large steps.

AdamW and weight decay

Pretraining usually uses AdamW, which applies weight decay (a pull of every weight towards zero, a form of regularization) separately from Adam's adaptive step. In Llama 3's scaling-law runs, the weight decay at each step was set to 0.1 times the learning rate at that step Established.

Gradient clipping

If the gradient's overall size (its norm) exceeds a threshold, scale it down to the threshold. One unusual batch then can't throw the weights far. OPT's team found that lowering the clipping threshold from 1.0 to 0.3 early in training helped stability Established.

What to remember

  • Warmup: increase the learning rate linearly from about zero over the first steps.
  • Cosine decay: follow half a cosine from the peak down to a small floor.
  • Llama 3 405B: peak 8 × 10⁻⁵, 8,000 warmup steps, cosine to 8 × 10⁻⁷ over 1.2M steps.
  • Bigger models use smaller peak learning rates (Llama 3: 3 × 10⁻⁴ for 8B, 8 × 10⁻⁵ for 405B).
  • Gradient clipping rescales the gradient if its norm exceeds a threshold.

Key papers

Important

Decoupled Weight Decay Regularization

Ilya Loshchilov, Frank Hutter · 2017 · ICLR 2019

Introduced AdamW, the variant of Adam used to train most modern Transformers and LLMs.

~40 min readarXiv:1711.05101✓ verified 2026-09-26
Optional

SGDR: Stochastic Gradient Descent with Warm Restarts

Ilya Loshchilov, Frank Hutter · 2016 · ICLR 2017

Introduced the cosine learning-rate schedule that most LLM pretraining runs still use, without the restarts.

How to read it: Equation 5 is the whole idea.

~20 min readarXiv:1608.03983✓ verified 2026-10-04
Optional

Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour

Priya Goyal, Piotr Dollár et al. · 2017

Showed how to train with very large batches across many GPUs without losing accuracy, using a linear learning-rate scaling rule and a gradual warmup.

~25 min readarXiv:1706.02677✓ verified 2026-10-04
Important

OPT: Open Pre-trained Transformer Language Models

Susan Zhang, Stephen Roller et al. · 2022

Released a GPT-3-sized model with its full training logbook: a rare, honest record of what goes wrong in a large run.

How to read it: Section 2.5, 'Training Processes', and the released logbook.

~30 min readarXiv:2205.01068✓ verified 2026-10-04
Essential

The Llama 3 Herd of Models

Aaron Grattafiori, Abhimanyu Dubey et al. · 2024

The most complete public account of building a frontier-scale model end to end: data pipeline, scaling-law experiments, 16,384-GPU training, failures and all.

How to read it: It is 90+ pages. For this chapter read Section 3 (pre-training) only: data, scaling laws, infrastructure and the training recipe.

~2 h readarXiv:2407.21783✓ verified 2026-10-04

Watch

4 h 1 min

Andrej Karpathy

Let's reproduce GPT-2 (124M)

Build and train the smallest GPT-2 from scratch in PyTorch, then optimise it. For when you want to implement what this chapter describes.

Frontier