Skip to content
Road to Intelligence

Concept · Chapter 11: Inside Modern LLMs

The KV Cache

Must knowKnow well14 minDifficulty

The KV cache stores every earlier token's keys and values in each attention layer, so generating a new token computes only its own query, key and value instead of reprocessing the whole sequence, at the cost of memory that grows with context length and with the number of users.

The problem

To generate token 1,001, attention needs keys and values for the 1,000 tokens before it. Recomputing them at every step repeats almost all of the previous step's work.

The solution

Keep each layer's keys and values for all tokens so far, append the new token's pair at every step, and let its query attend over the stored ones.

The consequence

Generation becomes cheap per token in arithmetic but memory-hungry: the cache grows linearly with context length and batch size, can exceed the model's own weights, and must be read in full at every decode step.

Why there's something to cache

In a decoder, token tt attends to tokens 1…t1 \dots t (causal masking). Its output depends on its own query and on the keys and values of every earlier token. Those earlier keys and values don't change when a new token arrives: token 7's key is computed from token 7's input, whatever comes later.

So compute them once and keep them. At each decode step:

  1. Compute the new token's query, key and value in every layer.
  2. Append the key and value to that layer's cache.
  3. Attend from the new query to all cached keys, and mix the cached values.

The prompt's keys and values are computed together during prefill; decoding then adds one entry per layer per token.

How big it gets

For each token, each layer stores one key and one value per key/value head, each a vector of the head dimension:

bytes per token=2×L×Hkv×dhead×bytes per number.\text{bytes per token} = 2 \times L \times H_{kv} \times d_{\text{head}} \times \text{bytes per number}.

Tiny example. The vLLM paper computes OPT-13B's cache as 2 (key and value) × 5,120 (hidden size) × 40 (layers) × 2 bytes = 800 KB per token Established. A 2,048-token request needs about 1.6 GB. That is modest next to the 26 GB of 16-bit weights, but it is per request: twenty such requests need more memory than the model.

Llama 2 7B uses one key/value head per query head: 2×32×32×128×22 \times 32 \times 32 \times 128 \times 2 bytes = 512 KiB per token. Serve 32 users with 4,096-token contexts and the cache is about 69 GB, five times the 14 GB of weights. The pattern scales: Pope and colleagues note that for a 500B+ model with multi-head attention, batch 512 and context 2,048, the KV cache totals 3 TB, three times the size of the parameters Established.

The bill arrives twice

The cache costs capacity: it must fit in GPU memory next to the weights, which caps the number of concurrent users and the context length. It also costs bandwidth: every decode step reads the whole cache of every sequence in the batch. Unlike the weights, it can't be shared between users: the KV cache is unique for each sequence in the batch Established, so batching doesn't amortize it.

That is why the next ideas in the chapter shrink or organize the cache: grouped-query attention stores fewer heads, quantization stores fewer bits, and paged attention stops wasting the memory reserved for it.

Try it

How Big Is the Cache?

Serve real Llama models to many users at once: add up weights and KV cache against GPU memory, change the number of key/value heads, and see the memory-bandwidth limit on tokens per second.

Know well8 min

Why should I care?

As a researcher

Long context is limited as much by cache memory and bandwidth as by attention's arithmetic, which is why so much architecture work (MQA, GQA, latent attention) changes the cache.

As an engineer

The cache, not the weights, often decides how many users fit on a GPU and how long a context you can afford. Its size is one multiplication you can do in your head.

Modern systems that depend on it

  • grouped-query attention
  • paged attention and vLLM
  • long-context serving
  • prompt caching

Historical context

Before

Each generated token re-ran attention over the full prefix, recomputing keys and values for every earlier token.

After

Keys and values are computed once per token and stored; memory per token per layer becomes a design target in its own right.

Used today

Every production LLM server keeps a KV cache, and serving systems are largely built around allocating, sharing and evicting it.

What to remember

  • Cache per token = 2 (K and V) × layers × KV heads × head dimension × bytes per number.
  • It grows linearly with context length and with the number of concurrent sequences.
  • OPT-13B: 800 KB per token, so one 2,048-token request needs about 1.6 GB.
  • Every decode step reads the whole cache for every sequence in the batch.
  • At long contexts and large batches the cache can be bigger than the weights.

Key papers

Essential

Efficiently Scaling Transformer Inference

Reiner Pope, Sholto Douglas et al. · 2022

The clearest account of why serving a large model is a memory problem: it splits inference into prefill and decode and shows when loading weights or the KV cache dominates.

How to read it: Read Section 2 (inference cost and the memory-time analysis). The worked example of a 3 TB KV cache is in Section 2.

~1 h readarXiv:2211.05102✓ verified 2026-10-05
Important

Fast Transformer Decoding: One Write-Head is All You Need

Noam Shazeer · 2019

Introduced multi-query attention: all heads share one key and one value, which shrinks the cache that every decoding step must reload.

How to read it: Short and readable. Sections 2.4 and 3.1 count memory accesses against arithmetic; Table 2 has the timings.

~25 min readarXiv:1911.02150✓ verified 2026-10-05
Essential

GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Joshua Ainslie, James Lee-Thorp et al. · 2023

Grouped-query attention is the compromise most open LLMs now use: a few key/value heads shared by groups of query heads.

How to read it: Figure 2 shows the three attention types side by side; Table 1 gives quality against inference time.

~25 min readarXiv:2305.13245✓ verified 2026-10-05
Important

DeepSeek-V3 Technical Report

DeepSeek-AI et al. · 2024

A detailed report on a frontier-scale open mixture-of-experts model that brings this chapter's ideas together: sparse experts, a compressed attention cache, FP8 training.

How to read it: Read Section 2 (architecture) first; Section 3 on infrastructure is valuable if you work on systems.

~1 h 30 min readarXiv:2412.19437✓ verified 2026-10-05
Important

Efficient Memory Management for Large Language Model Serving with PagedAttention

Woosuk Kwon, Zhuohan Li et al. · 2023

Brought operating-system paging to the KV cache (the vLLM server), so many more requests fit in GPU memory at once.

How to read it: Section 3 (why existing systems waste memory) and Figure 6 (the block table) are the core.

~45 min readarXiv:2309.06180✓ verified 2026-10-05

Watch