Concept · Chapter 11: Inside Modern LLMs
The KV Cache
The KV cache stores every earlier token's keys and values in each attention layer, so generating a new token computes only its own query, key and value instead of reprocessing the whole sequence, at the cost of memory that grows with context length and with the number of users.
The problem
To generate token 1,001, attention needs keys and values for the 1,000 tokens before it. Recomputing them at every step repeats almost all of the previous step's work.
The solution
Keep each layer's keys and values for all tokens so far, append the new token's pair at every step, and let its query attend over the stored ones.
The consequence
Generation becomes cheap per token in arithmetic but memory-hungry: the cache grows linearly with context length and batch size, can exceed the model's own weights, and must be read in full at every decode step.
You should understand first
- Vectors
- Dot Product
- Embeddings
- Attention
- Probability and Distributions
- Softmax
- Self-Attention
- Causal Masking
- Text as Data
- Conditional Probability and Bayes' Theorem
- Probability of Sequences
- Language Modeling
- Entropy
- Loss Functions
- Cross-Entropy Loss
- Autoregressive Next-Token Prediction
- Matrix Multiplication
- Multi-Head Attention
- Positional Encoding
- Residual Connections
- Layer Normalization
- Feed-Forward Sublayer (MLP)
- The Transformer Block
- Prefill, Decode and the Memory Wall
- The KV Cache
Why there's something to cache
In a decoder, token attends to tokens (causal masking). Its output depends on its own query and on the keys and values of every earlier token. Those earlier keys and values don't change when a new token arrives: token 7's key is computed from token 7's input, whatever comes later.
So compute them once and keep them. At each decode step:
- Compute the new token's query, key and value in every layer.
- Append the key and value to that layer's cache.
- Attend from the new query to all cached keys, and mix the cached values.
The prompt's keys and values are computed together during prefill; decoding then adds one entry per layer per token.
How big it gets
For each token, each layer stores one key and one value per key/value head, each a vector of the head dimension:
Tiny example. The vLLM paper computes OPT-13B's cache as 2 (key and value) × 5,120 (hidden size) × 40 (layers) × 2 bytes = 800 KB per token Established. A 2,048-token request needs about 1.6 GB. That is modest next to the 26 GB of 16-bit weights, but it is per request: twenty such requests need more memory than the model.
Llama 2 7B uses one key/value head per query head: bytes = 512 KiB per token. Serve 32 users with 4,096-token contexts and the cache is about 69 GB, five times the 14 GB of weights. The pattern scales: Pope and colleagues note that for a 500B+ model with multi-head attention, batch 512 and context 2,048, the KV cache totals 3 TB, three times the size of the parameters Established.
The bill arrives twice
The cache costs capacity: it must fit in GPU memory next to the weights, which caps the number of concurrent users and the context length. It also costs bandwidth: every decode step reads the whole cache of every sequence in the batch. Unlike the weights, it can't be shared between users: the KV cache is unique for each sequence in the batch Established, so batching doesn't amortize it.
That is why the next ideas in the chapter shrink or organize the cache: grouped-query attention stores fewer heads, quantization stores fewer bits, and paged attention stops wasting the memory reserved for it.
Try it
Serve real Llama models to many users at once: add up weights and KV cache against GPU memory, change the number of key/value heads, and see the memory-bandwidth limit on tokens per second.
Why should I care?
As a researcher
Long context is limited as much by cache memory and bandwidth as by attention's arithmetic, which is why so much architecture work (MQA, GQA, latent attention) changes the cache.
As an engineer
The cache, not the weights, often decides how many users fit on a GPU and how long a context you can afford. Its size is one multiplication you can do in your head.
Modern systems that depend on it
- grouped-query attention
- paged attention and vLLM
- long-context serving
- prompt caching
Historical context
Before
Each generated token re-ran attention over the full prefix, recomputing keys and values for every earlier token.
After
Keys and values are computed once per token and stored; memory per token per layer becomes a design target in its own right.
Used today
Every production LLM server keeps a KV cache, and serving systems are largely built around allocating, sharing and evicting it.
What to remember
- Cache per token = 2 (K and V) × layers × KV heads × head dimension × bytes per number.
- It grows linearly with context length and with the number of concurrent sequences.
- OPT-13B: 800 KB per token, so one 2,048-token request needs about 1.6 GB.
- Every decode step reads the whole cache for every sequence in the batch.
- At long contexts and large batches the cache can be bigger than the weights.
Key papers
Efficiently Scaling Transformer Inference
Reiner Pope, Sholto Douglas et al. · 2022
The clearest account of why serving a large model is a memory problem: it splits inference into prefill and decode and shows when loading weights or the KV cache dominates.
How to read it: Read Section 2 (inference cost and the memory-time analysis). The worked example of a 3 TB KV cache is in Section 2.
Fast Transformer Decoding: One Write-Head is All You Need
Noam Shazeer · 2019
Introduced multi-query attention: all heads share one key and one value, which shrinks the cache that every decoding step must reload.
How to read it: Short and readable. Sections 2.4 and 3.1 count memory accesses against arithmetic; Table 2 has the timings.
GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints
Joshua Ainslie, James Lee-Thorp et al. · 2023
Grouped-query attention is the compromise most open LLMs now use: a few key/value heads shared by groups of query heads.
How to read it: Figure 2 shows the three attention types side by side; Table 1 gives quality against inference time.
DeepSeek-V3 Technical Report
DeepSeek-AI et al. · 2024
A detailed report on a frontier-scale open mixture-of-experts model that brings this chapter's ideas together: sparse experts, a compressed attention cache, FP8 training.
How to read it: Read Section 2 (architecture) first; Section 3 on infrastructure is valuable if you work on systems.
Efficient Memory Management for Large Language Model Serving with PagedAttention
Woosuk Kwon, Zhuohan Li et al. · 2023
Brought operating-system paging to the KV cache (the vLLM server), so many more requests fit in GPU memory at once.
How to read it: Section 3 (why existing systems waste memory) and Figure 6 (the block table) are the core.
Watch
Stanford Online
Stanford CS336 Language Modeling from Scratch | Spring 2025 | Lecture 10: Inference
The closest single lecture to this chapter: the arithmetic of serving a language model and the main ways to make it cheaper.