Skip to content
Road to Intelligence

Concept · Chapter 11: Inside Modern LLMs

Prefill, Decode and the Memory Wall

Must knowKnow well14 minDifficulty

Serving a language model has two phases: prefill reads the whole prompt in one parallel pass and keeps the GPU busy, while decode produces one token per pass and spends most of its time moving weights and cached keys and values out of memory.

The problem

A GPU that trains a model at a good fraction of its peak speed sits mostly idle when the same model generates text. Why?

The solution

Count bytes as well as FLOPs. Each decode step must read every weight (and every cached key and value) to do only about one multiply-add per weight per sequence, so the time is set by memory bandwidth; batching many users shares those reads.

The consequence

Almost every inference technique in this chapter attacks the bytes a decode step reads: smaller caches (GQA, paging), smaller weights (quantization), fewer weights used (mixture of experts) or more tokens per read (batching, speculative decoding).

Two phases with different bottlenecks

Pope and colleagues split the latency of generating text into the time to process the prompt, which they call prefill, and the time to generate output tokens one at a time, decode Established.

Prefill looks like training: all prompt tokens are known, so they go through each layer together as one large matrix multiplication. The GPU's arithmetic units are busy. This phase decides the time to first token.

Decode is different. Each step produces one new token per sequence, and the next step can't start until it exists (autoregressive generation). A step is a forward pass for a single position, but it still has to read every weight matrix from GPU memory.

Count the bytes, not just the FLOPs

A useful ratio is arithmetic intensity: floating-point operations done per byte moved from memory.

Take one weight matrix with 16-bit (2-byte) entries. In a decode step, each weight is used for one multiply-add (2 FLOPs) per sequence in the batch BB, and is read once:

intensity=2B FLOPs per weight2 bytes per weight=B.\text{intensity} = \frac{2B \text{ FLOPs per weight}}{2 \text{ bytes per weight}} = B.

With one user, that is 1 FLOP per byte. An H100 SXM reads memory at 3.35 TB/s Established and does about 989 trillion dense BF16 operations per second, so it needs about 295 FLOPs per byte to keep its arithmetic busy; Stanford's CS336 inference lecture concludes that such a matrix multiplication is compute-limited only when the batch exceeds about 295 Established. Below that, the step waits on memory. (NVIDIA's headline figure of 1,979 BF16 teraFLOPS is quoted "with sparsity"; the dense rate is half that.)

A tiny calculation

Llama 3 8B has about 8 billion parameters: 16 GB in 16 bits. Each decode step reads all 16 GB.

time per step≥16 GB3.35 TB/s≈4.8 ms⇒at most about 210 tokens per second.\text{time per step} \ge \frac{16 \text{ GB}}{3.35 \text{ TB/s}} \approx 4.8 \text{ ms} \quad\Rightarrow\quad \text{at most about 210 tokens per second}.

That ceiling holds no matter how fast the GPU's arithmetic is. Now serve 32 users at once. The weights are read once per step for all 32, so total throughput rises roughly 32-fold, until the other thing every step reads, each user's KV cache, takes over. Pope and colleagues found that at small batch sizes and short sequences loading the weights dominates, while at large batches and long contexts loading the KV cache dominates Established.

What it explains

Read the rest of the chapter as a list of ways to cut the bytes per generated token:

TechniqueWhat it shrinks
Grouped-query attentionthe KV cache read per token
Quantizationthe bytes per weight
Mixture of expertsthe weights touched per token
Batching, paged attentionweight reads per user, wasted cache memory
Speculative decodingtarget-model passes per token

Try it

How Big Is the Cache?

Serve real Llama models to many users at once: add up weights and KV cache against GPU memory, change the number of key/value heads, and see the memory-bandwidth limit on tokens per second.

Know well8 min

Why should I care?

As a researcher

Inference cost now shapes architecture: grouped-query attention, mixture-of-experts and long-context methods are judged by what they do to memory traffic as much as by quality.

As an engineer

Latency, throughput and cost per token all follow from a short calculation: bytes read per step divided by memory bandwidth. It tells you when batching helps and why a bigger GPU doesn't always mean faster replies.

Modern systems that depend on it

  • the KV cache
  • grouped-query attention
  • quantization
  • speculative decoding
  • batching and paged attention

Historical context

Before

Inference was treated as the easy part: one forward pass per token, cheap compared with training.

After

Serving is planned like training: a memory budget (weights plus caches), a bandwidth budget per token, and techniques chosen by which of the two they relieve.

Used today

Every LLM serving system separates prefill from decode, batches decode steps across users, and reports time to first token (prefill) separately from tokens per second (decode).

What to remember

  • Prefill: the whole prompt in one parallel pass, compute-bound like training.
  • Decode: one token per pass; every weight and cached key/value is read once per step.
  • Arithmetic intensity = FLOPs ÷ bytes moved. Decoding one sequence with 16-bit weights: about 1.
  • An H100 does about 295 dense BF16 FLOPs per byte it reads, so batch-1 decoding leaves most of its arithmetic idle.
  • Lower bound on a decode step: bytes read ÷ memory bandwidth. Batching shares the weight reads across users.

Key papers

Essential

Efficiently Scaling Transformer Inference

Reiner Pope, Sholto Douglas et al. · 2022

The clearest account of why serving a large model is a memory problem: it splits inference into prefill and decode and shows when loading weights or the KV cache dominates.

How to read it: Read Section 2 (inference cost and the memory-time analysis). The worked example of a 3 TB KV cache is in Section 2.

~1 h readarXiv:2211.05102✓ verified 2026-10-05
Important

Fast Transformer Decoding: One Write-Head is All You Need

Noam Shazeer · 2019

Introduced multi-query attention: all heads share one key and one value, which shrinks the cache that every decoding step must reload.

How to read it: Short and readable. Sections 2.4 and 3.1 count memory accesses against arithmetic; Table 2 has the timings.

~25 min readarXiv:1911.02150✓ verified 2026-10-05

Watch