Skip to content
Road to Intelligence

Concept · Chapter 9: How an LLM Is Actually Built

The Compute Budget (C ≈ 6ND)

Must knowKnow well12 minDifficulty

Training a dense Transformer with N parameters on D tokens costs about 6ND floating-point operations, which, divided by what a GPU actually sustains, turns a model plan into GPU-days and money.

The problem

Before spending months and millions on a run, a team needs to know how much computation it requires and how long it will take on the hardware it has.

The solution

Count FLOPs per token (about 2N for the forward pass and 4N for the backward pass), multiply by tokens, and divide by the throughput each GPU really achieves (its model FLOPs utilisation times its peak).

The consequence

Compute became the common currency of AI progress: budgets, scaling laws, policy thresholds and comparisons between models are all stated in training FLOPs.

Where 6 comes from

Most of a Transformer's work is multiplying activations by weight matrices. Each weight is used once per token in the forward pass, as one multiply and one add: about 2N FLOPs per token for N parameters. The backward pass computes gradients with respect to both the activations and the weights, about twice as much again: 4N. Kaplan and colleagues therefore estimated training compute as C ≈ 6N floating-point operations per training token Established:

C≈6 N DC \approx 6\,N\,D

The estimate leaves out attention's own cost, which grows with the context length; it is small whenever the model's width is large relative to the context Established.

Checking it against real runs

ModelND6NDReported
GPT-3175B300B3.15 × 10²³3.14 × 10²³ Established
Chinchilla70B1.4T5.9 × 10²³5.76 × 10²³ (the Gopher budget) Established
Llama 3 405B405B15.6T3.79 × 10²⁵3.8 × 10²⁵ Established

From FLOPs to days

GPUs never run at their peak. Model FLOPs utilisation (MFU) is the useful model FLOPs per second divided by the hardware's peak. PaLM 540B reached 46.2% MFU Established; Llama 3 405B reached 38–43%, about 400 TFLOP/s per H100 on 16,384 GPUs Established.

Tiny example. 3.8 × 10²⁵ FLOPs ÷ (16,384 GPUs × 4 × 10¹⁴ FLOP/s) ≈ 5.8 × 10⁶ seconds ≈ 67 days of uninterrupted training at that rate. This is our arithmetic, not a reported duration; real runs also lose time to failures and restarts.

Why should I care?

As a researcher

Every scaling-law plot has compute on its x-axis; a claim about efficiency is a claim about loss per FLOP.

As an engineer

6ND plus a utilisation figure is the napkin maths for whether a training or fine-tuning job is hours, days or impossible on your hardware.

Modern systems that depend on it

  • scaling laws
  • compute-optimal training
  • cluster sizing
  • cost estimates

Historical context

Before

Training cost was described by GPU type and wall-clock time, which made runs hard to compare.

After

Runs are described by total FLOPs and achieved utilisation; 6ND is the standard first estimate.

Used today

Model reports quote training FLOPs (Llama 3 405B: 3.8 × 10²⁵); MFU is the standard efficiency metric for large runs.

What to remember

  • C ≈ 6ND: about 2N FLOPs per token forward, 4N backward.
  • GPT-3: 6 × 175B × 300B ≈ 3.1 × 10²³ FLOPs. Llama 3 405B: 6 × 405B × 15.6T ≈ 3.8 × 10²⁵.
  • MFU = achieved model FLOPs per second ÷ hardware peak. Large runs reach roughly 40–50%.
  • Days ≈ C ÷ (GPUs × FLOP/s per GPU × 86,400).
  • 6ND ignores attention's sequence-length term, which matters only for long contexts relative to model width.

Key papers

Essential

Language Models are Few-Shot Learners

Tom B. Brown, Benjamin Mann et al. · 2020 · NeurIPS 2020

GPT-3 (175B parameters) showed that a large enough language model can perform new tasks from a few examples in its prompt, without any gradient updates.

How to read it: 75 pages. Sections 1–2 and Figure 1.2 carry the core idea; Section 6 on broader impacts is worth reading too.

~1 h 30 min readarXiv:2005.14165✓ verified 2026-09-26
Essential

Scaling Laws for Neural Language Models

Jared Kaplan, Sam McCandlish et al. · 2020

Found that language-model loss falls as a smooth power law in parameters, data and compute — making model scale something you could plan.

~1 h readarXiv:2001.08361✓ verified 2026-09-26
Important

PaLM: Scaling Language Modeling with Pathways

Aakanksha Chowdhery, Sharan Narang et al. · 2022

A 540-billion-parameter model whose report defined model FLOPs utilization (MFU) and described, unusually frankly, the loss spikes of a very large run.

How to read it: Sections 4 (training infrastructure) and 5.1 (training instability) are the parts for this chapter.

~1 h readarXiv:2204.02311✓ verified 2026-10-04
Essential

The Llama 3 Herd of Models

Aaron Grattafiori, Abhimanyu Dubey et al. · 2024

The most complete public account of building a frontier-scale model end to end: data pipeline, scaling-law experiments, 16,384-GPU training, failures and all.

How to read it: It is 90+ pages. For this chapter read Section 3 (pre-training) only: data, scaling laws, infrastructure and the training recipe.

~2 h readarXiv:2407.21783✓ verified 2026-10-04

Watch