Concept · Chapter 11: Inside Modern LLMs
Multi-Query and Grouped-Query Attention
Multi-query and grouped-query attention keep many query heads but share fewer key/value heads between them (one in MQA, a few groups in GQA), shrinking the KV cache and the memory each decode step must read, with little loss in quality.
The problem
With multi-head attention, every head keeps its own keys and values in the cache, so the cache, and the bytes read per decode step, grow with the number of heads.
The solution
Let groups of query heads share one key head and one value head. The queries still differ per head, so heads can attend to different things; only the stored keys and values are shared.
The consequence
The cache shrinks by the ratio of query heads to key/value heads (4× to 16× in Llama 3). GQA with a handful of KV heads became the default in open LLMs; later designs compress the cache further.
You should understand first
- Vectors
- Dot Product
- Embeddings
- Attention
- Probability and Distributions
- Softmax
- Self-Attention
- Multi-Head Attention
- Causal Masking
- Text as Data
- Conditional Probability and Bayes' Theorem
- Probability of Sequences
- Language Modeling
- Entropy
- Loss Functions
- Cross-Entropy Loss
- Autoregressive Next-Token Prediction
- Matrix Multiplication
- Positional Encoding
- Residual Connections
- Layer Normalization
- Feed-Forward Sublayer (MLP)
- The Transformer Block
- Prefill, Decode and the Memory Wall
- The KV Cache
- Multi-Query and Grouped-Query Attention
Share the keys, keep the questions
In multi-head attention, each of heads has its own query, key and value projections. During decoding, all the keys and values go into the KV cache and are reread at every step.
Shazeer's multi-query attention keeps the separate query heads but shares a single key head and value head across all of them, because reloading the large key and value tensors made incremental decoding memory-bound Established. Each head still asks its own question (its own query) over the same stored keys and values.
Grouped-query attention generalizes this: the query heads are split into groups, and each group shares one key head and one value head, so the number of KV heads lies between one (MQA) and the number of query heads (MHA) Established.
| Query heads | Key/value heads | Cache per token, per layer | |
|---|---|---|---|
| Multi-head (MHA) | numbers | ||
| Grouped-query (GQA) | |||
| Multi-query (MQA) | 1 |
What it bought
In Shazeer's translation experiments, the decoder's time per token fell from 46 µs with multi-head attention to 3.8 µs with multi-query attention, and quality was only slightly worse Established.
Quality loss and the cost of retraining kept many models on MHA. Ainslie and colleagues converted existing multi-head checkpoints with about 5% of the original pretraining compute ("uptraining"). On their T5-XXL benchmarks, GQA with 8 groups averaged 47.1 against 47.2 for multi-head attention and 46.6 for MQA, at an inference time of 0.28 s against 1.51 s for multi-head Established.
Llama 2 adopted grouped-query attention for its larger models Established, and Llama 3 uses 8 key/value heads at 8B, 70B and 405B parameters Established.
Tiny example. Llama 3 405B: 126 layers, 128 query heads, 8 KV heads of dimension 128. Cache per token in 16 bits: bytes ≈ 504 KiB. With one KV head per query head it would be 16 times that, about 7.9 MiB per token. Llama 2 7B, with full multi-head attention, stores 512 KiB per token: a model 58 times smaller with a slightly larger cache per token.
Further compression
GQA isn't the end of the line. DeepSeek-V3's multi-head latent attention jointly compresses keys and values into a low-rank latent vector to reduce the KV cache during inference Established. The direction is the same: store less per token, read less per step.
What to remember
- MHA: one K/V head per query head. MQA: one K/V head for all. GQA: one per group.
- Cache shrinks by (query heads) ÷ (KV heads); queries are not shared.
- MQA cut a translation decoder's time per token from 46 µs to 3.8 µs, with a small quality loss.
- GQA-8 was about as good as multi-head attention and about as fast as MQA; converting a checkpoint took ~5% of pretraining compute.
- Llama 3 uses 8 KV heads at every size: 405B stores less cache per token than Llama 2 7B.
Key papers
Efficiently Scaling Transformer Inference
Reiner Pope, Sholto Douglas et al. · 2022
The clearest account of why serving a large model is a memory problem: it splits inference into prefill and decode and shows when loading weights or the KV cache dominates.
How to read it: Read Section 2 (inference cost and the memory-time analysis). The worked example of a 3 TB KV cache is in Section 2.
Fast Transformer Decoding: One Write-Head is All You Need
Noam Shazeer · 2019
Introduced multi-query attention: all heads share one key and one value, which shrinks the cache that every decoding step must reload.
How to read it: Short and readable. Sections 2.4 and 3.1 count memory accesses against arithmetic; Table 2 has the timings.
GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints
Joshua Ainslie, James Lee-Thorp et al. · 2023
Grouped-query attention is the compromise most open LLMs now use: a few key/value heads shared by groups of query heads.
How to read it: Figure 2 shows the three attention types side by side; Table 1 gives quality against inference time.
DeepSeek-V3 Technical Report
DeepSeek-AI et al. · 2024
A detailed report on a frontier-scale open mixture-of-experts model that brings this chapter's ideas together: sparse experts, a compressed attention cache, FP8 training.
How to read it: Read Section 2 (architecture) first; Section 3 on infrastructure is valuable if you work on systems.
Llama 2: Open Foundation and Fine-Tuned Chat Models
Hugo Touvron, Louis Martin et al. · 2023
A widely used open model family; its larger models adopted grouped-query attention for faster inference.
How to read it: For this chapter, read Section 2.2 and Appendix A.2.1, which compares multi-head, multi-query and grouped-query attention.
Watch
Stanford Online
Stanford CS336 Lang. Modeling from Scratch | Spring 2025 | Lec. 3: Architectures, Hyperparameters
A survey of what changed inside the Transformer block between 2017 and today's open models, and which choices nearly everyone now agrees on.
Stanford Online
Stanford CS336 Language Modeling from Scratch | Spring 2025 | Lecture 10: Inference
The closest single lecture to this chapter: the arithmetic of serving a language model and the main ways to make it cheaper.