Part III · Large Language Models
Chapter 11
Inside Modern LLMs
The engineering that makes large models fast, cheap and adaptable.
In one sentenceKV caches, efficient attention, mixture-of-experts, quantization and low-rank adaptation each trade something away to fit large models onto real hardware.
The problem
The serving bill
Chapter 9 trained Llama 3 405B. Chapter 10 turned a checkpoint like it into an assistant. Now someone has to run it, for many people at once, quickly and at a price they'll pay. Training happens once; inference happens every time anyone asks anything.
Meta's Llama 3 report describes this part too, and it makes a good running example:
| What | Llama 3 405B | Section |
|---|---|---|
| Weights in 16 bits | 810 GB: more than one server's eight 80 GB GPUs | The serving bill |
| Cache per token of context | about 0.5 MB, thanks to grouped-query attention | KV cache; Fewer keys and values |
| Longest context | 128K tokens, about 68 GB of cache for one such conversation | The quadratic wall |
| Served on | 16 GPUs across two servers in BF16 | The serving bill |
| FP8 weights and activations in the feed-forward layers | up to 50% more prefill throughput | Fewer bits per weight |
The report states that in BF16 the 405B model does not fit in the memory of a single machine with 8 H100 GPUs, so Meta ran inference across 16 GPUs on two machines, and that FP8 quantization improved prefill throughput by up to 50% Established. The 810 GB, 0.5 MB and 68 GB are our arithmetic from the published architecture, explained below.
Inference has two phases. Prefill reads the prompt in one parallel pass, much like training. Decode then produces one token per pass, and every pass must read all the weights from memory. An H100 reads memory at 3.35 TB/s Established and can do about 295 dense BF16 operations in the time it reads one byte. A decode step for one user does about one operation per byte of weights. So during decoding the GPU's arithmetic units mostly wait: the speed limit is memory bandwidth, not FLOPs.
That observation organizes the whole chapter. Each idea below reduces the bytes moved per generated token, or the memory that limits how many users share a pass.
The first saving
Remember, don't recompute
To generate the next token, attention needs keys and values for every token before it. Those don't change as the text grows, so the model computes them once and stores them: the KV cache. Each new token then needs only its own query, key and value.
The cache isn't small. Per token it holds a key and a value for every layer and every key/value head. For OPT-13B that is 800 KB per token, so one 2,048-token request needs up to 1.6 GB Established. Unlike the weights, each conversation has its own cache, and every decode step reads all of it.
Try it below with real model configurations. With Llama 3 8B and one user at 8,192 tokens, the 16 GB of weights dominate, and memory bandwidth caps decoding at about 196 tokens per second. Raise the users to 32: each gets about 67 tokens per second, but the GPU produces over 2,100 in total, because the weights are read once per step for everyone. Now switch the attention to "one per query head": the cache grows to 137 GB and nothing fits.
Try it
Serve real Llama models to many users at once: add up weights and KV cache against GPU memory, change the number of key/value heads, and see the memory-bandwidth limit on tokens per second.
The second saving
Fewer keys and values
The cache grows with the number of key/value heads, and the fix is to have fewer of them. Multi-query attention keeps every query head but shares one key head and one value head among them; in its 2019 translation experiments, decoder time per token fell from 46 µs to 3.8 µs with only minor quality loss Established. Grouped-query attention sits in between: groups of query heads share a key/value head. Converted from multi-head checkpoints with about 5% of the original pretraining compute, GQA with 8 groups came close to multi-head quality at close to multi-query speed Established.
Llama 3 uses 8 key/value heads at every size Established. The result is striking: the 405B model stores about 504 KiB of cache per token, slightly less than Llama 2 7B's 512 KiB with full multi-head attention, despite having 58 times as many parameters.
Long contexts
The quadratic wall
Chapter 7 noted that attention compares every token with every other: tokens give an score matrix per head. At 8,192 tokens that is 67 million scores per head per layer. Double the context and the scores quadruple.
For years the standard implementation wrote that whole matrix to the GPU's main memory, read it back for the softmax, and read it again to multiply by the values. FlashAttention never writes it. It processes attention in tiles small enough for the GPU's fast on-chip memory, keeping a running maximum and sum so the softmax can be completed tile by tile. The output is exactly the same. It trained GPT-2 3× faster at 1K tokens and made attention's memory linear in sequence length Established. The arithmetic was never the bottleneck; the trips to memory were.
Small changes that stuck
The modern block
Open the code of a recent open model and you'll recognize Chapter 7's block, with a few edits. Llama 2 describes its architecture as the standard Transformer with pre-normalization using RMSNorm, the SwiGLU activation and rotary positional embeddings Established:
- Pre-norm: normalize at the start of each sublayer, not after adding the residual, so the residual path is a clean sum and training is less fragile.
- RMSNorm: divide by the root mean square and skip LayerNorm's mean subtraction; cheaper, and in practice just as good.
- SwiGLU: a gated feed-forward layer, one projection multiplied by an activated second one. Its author found it improved quality and offered no explanation for why Established.
The fourth edit is position. Instead of adding a position vector to each token, RoPE rotates queries and keys by angles proportional to their positions, so attention scores depend on how far apart two tokens are. Its frequencies are also the main dial for stretching a model's context beyond what it was trained on.
Parameters without the compute
Many experts, few at a time
A dense model uses every parameter for every token. A mixture-of-experts layer replaces the feed-forward network with many smaller experts and a router that sends each token to a few of them. Memory holds all the experts; each token pays for only the ones it uses.
Mixtral 8x7B routes each token to 2 of 8 experts per layer: 47B parameters in total, 13B used per token. It matched or beat Llama 2 70B and GPT-3.5 on the benchmarks its authors report Established. DeepSeek-V3 has 671B parameters, of which 37B are activated for each token Established. Step through one token's trip through a layer:
- Expert 11.2idle
- Expert 2-0.3idle
- Expert 30.52runs
- Expert 40.1idle
- Expert 50.4idle
- Expert 6-1.0idle
- Expert 70.48runs
- Expert 80.0idle
Output = 0.52 × expert 3 + 0.48 × expert 7. Six experts do no work for this token, but all eight stay in memory. Try another token: it can choose different experts.
Mixtral 8x7B, per token
Router scores are made up for illustration; the top-2 selection and softmax are computed. Mixtral’s parameter counts are from its paper.
The savings are in arithmetic, not memory: Mixtral's authors note that its serving memory is proportional to its 47B total parameters Established. MoE also brings its own engineering: keeping the experts evenly used, and moving tokens between the GPUs that hold different experts.
Smaller numbers
Fewer bits per weight
If decoding is limited by bytes, store fewer bytes per weight. Quantization replaces each 16-bit weight by an 8- or 4-bit integer and a scale shared by a group of weights. Llama 3 70B's weights take 140 GB in 16 bits and about 39 GB at 4 bits, small enough for a single 80 GB GPU.
The difficulty is a few very large values. A shared scale has to stretch to cover the largest number in its group, and everything else gets squeezed into a few levels. Dettmers and colleagues found that from about 6.7B parameters, large language models develop systematic outlier features that ruin naive 8-bit quantization Established, and handled them by keeping those few dimensions in 16 bits.
See it on a toy row of 512 weights below. In 8 bits with one scale for the row, the error is under 1%. Drop to 4 bits: about 11%. Add one outlier: at 4 bits, 99.6% of the weights now round to zero. Switch to one scale per 64 weights and the damage is confined to the outlier's own group, at a cost of half a bit per weight for the extra scales.
Try it · toy model
Round a row of weights to 8, 4, 3 or 2 bits, add one outlier, and see why a shared scale fails and per-group scales rescue it.
Better methods round more cleverly. GPTQ quantized 175B-parameter models to 3–4 bits in about four GPU hours by adjusting the remaining weights to compensate for each rounding error Established, and AWQ protects the roughly 1% of weights that matter most for the activations they meet.
Compression
Smaller models from bigger ones
Two older ideas make models smaller in different ways. Distillation trains a small student to match a large teacher's output probabilities, which say more than the right answer alone: that a 2 looks a bit like a 7. DistilBERT was 40% smaller and 60% faster than BERT while keeping 97% of its language-understanding performance Established.
Pruning removes weights that contribute little. Trained networks tolerate a lot of it, but zeros only save time when the hardware can skip them, so patterns matter more than counts.
Adapting cheaply
Adapting without retraining
Fine-tuning every weight of a large model, as in Chapter 10, needs memory for gradients and optimizer states on every parameter, and produces a full copy of the model per task. LoRA freezes the model and learns, for selected matrices, a low-rank update: two thin matrices whose product is added to the weights. Slide the rank and watch the number of trained parameters:
- This matrix, full fine-tuning
- 16,777,216
- This matrix, LoRA
- 65,536 (0.39%)
- Llama 3 8B, query + value in all 32 layers
- 3.4M (0.043% of 8B)
Trainable numbers are r × (d + k) per matrix. B starts at zero, so training begins from the unchanged model. Llama 3 8B’s value projection is 4,096 × 1,024 because of grouped-query attention. Counts are our arithmetic from the published architecture; Adam’s states are needed only for the trainable part.
On GPT-3 175B, LoRA cut trainable parameters 10,000-fold and GPU memory 3-fold compared with full fine-tuning with Adam, while matching its quality, and the update can be merged into the weights so inference is no slower Established. QLoRA keeps the frozen model in 4 bits and fine-tuned a 65B-parameter model on a single 48 GB GPU Established.
Using the idle arithmetic
Guess ahead, check in parallel
Decoding leaves arithmetic idle, and checking five tokens costs about as much as generating one. Speculative decoding puts that to use: a small draft model guesses several tokens ahead, and the large model scores all of them in one pass, keeping each guess by a rule that, remarkably, leaves the output distribution exactly the large model's own.
In the lab, the target model prefers "the" (0.5) and the draft prefers "a" (0.4). Run thousands of draft-and-verify steps: the words come out at the target's proportions, not the draft's, and the draft was accepted 80% of the time. Then set the acceptance rate, draft length and draft cost to see the expected speed-up.
Try it · toy model
Run speculative decoding's accept-or-resample rule thousands of times and watch the output keep the large model's distribution, then trade acceptance rate, draft length and draft cost for speed.
Speculative decoding ran T5-XXL 2–3× faster with identical outputs Established.
The server
Serving many users
Batching is the most powerful lever of all, because each pass's weight reads are shared by every request in the batch. What limits the batch is cache memory. The vLLM authors measured that existing servers used only 20.4–38.2% of their KV-cache memory for actual tokens Established; the rest was reserved for replies that might grow, or lost to fragmentation.
PagedAttention borrows virtual memory from operating systems: the cache lives in small blocks allocated as tokens arrive and found through a table, and requests with a common prompt share blocks. Together with batching at the level of individual decode steps, so finished requests leave and new ones join at every step, the vLLM server improved throughput 2–4× at the same latency over the systems it was compared with Established.
The bigger picture
Why it matters
None of this changes what a language model is. The model of Chapter 8 still predicts the next token; the block of Chapter 7 is still recognizably there. What changed is the cost of each token, and cost decides what gets built: longer contexts, many samples per question, models on laptops and phones, fine-tuned variants for every customer.
For an engineer, the habit to take away is the napkin calculation: bytes per token, bytes per second, and which of weights, cache or communication dominates. It predicts which technique will help before you try it. For a researcher, the lesson is that architecture now answers to hardware. Grouped-query attention, mixture-of-experts and FlashAttention are judged as much by memory traffic as by loss. Inference cost is increasingly a design constraint at training time too, as Chapter 9's "train smaller, longer" argument showed Interpretation.
Each technique trades something: quantization and distillation give up a little quality; mixture-of-experts gives up memory and simplicity; speculative decoding spends arithmetic; LoRA restricts what an update can express. None is free, and published gains are measured on particular models and tasks.
What comes next
What comes next
A served model is fast and cheap, but its knowledge is still frozen at the end of training, and it can't say where an answer came from. Chapter 12 connects a model to documents it has never seen: embeddings that measure meaning, retrieval that finds the relevant passages, and the application stack around them. The KV cache returns there too, as the cost of every retrieved passage you put in the context.
From a trained checkpoint to a served model
Concepts in this chapter
Mark each one as you go. Must-know concepts are the core path.
- The Quadratic Cost of AttentionSelf-attention compares every token with every other token, so its score matrix, and a naive implementation's time and memory, grow with the square of the sequence length.UnderstandMust know
- Multi-Query and Grouped-Query AttentionMulti-query and grouped-query attention keep many query heads but share fewer key/value heads between them (one in MQA, a few groups in GQA), shrinking the KV cache and the memory each decode step must read, with little loss in quality.UnderstandMust know
- The KV CacheThe KV cache stores every earlier token's keys and values in each attention layer, so generating a new token computes only its own query, key and value instead of reprocessing the whole sequence, at the cost of memory that grows with context length and with the number of users.Know wellMust know
- LoRA, QLoRA and Parameter-Efficient Fine-TuningLoRA fine-tunes a large model by freezing its weights and learning, for chosen weight matrices, a small low-rank update (the product of two thin matrices), so only a tiny fraction of parameters is trained and stored per task.Know wellMust know
- Mixture of Experts (MoE)A mixture-of-experts layer replaces one feed-forward network with many smaller 'experts' and a router that sends each token to only a few of them, so the model can store far more parameters than it uses for any one token.Know wellMust know
- Prefill, Decode and the Memory WallServing a language model has two phases: prefill reads the whole prompt in one parallel pass and keeps the GPU busy, while decode produces one token per pass and spends most of its time moving weights and cached keys and values out of memory.Know wellMust know
- QuantizationQuantization stores weights (and sometimes activations or the KV cache) in fewer bits, such as 8 or 4 instead of 16, by mapping each group of numbers onto a small grid of levels with a shared scale, trading a little accuracy for much less memory and memory traffic.Know wellMust know
- FlashAttentionFlashAttention computes exact attention in tiles that stay in the GPU's small, fast on-chip memory, combining partial softmaxes as it goes, so the n × n score matrix is never written to slow memory: same result, much less memory traffic, memory linear in sequence length.UnderstandShould know
- Knowledge DistillationKnowledge distillation trains a small 'student' model to match a large 'teacher' model's output probabilities, which carry more information than the correct answers alone, so the student gets closer to the teacher than training on labels would allow.UnderstandShould know
- The Modern Decoder Block: Pre-Norm, RMSNorm, SwiGLUToday's open LLMs keep the 2017 Transformer's decoder block but normalize before each sublayer instead of after (pre-norm), use the cheaper RMSNorm, and replace the feed-forward layer's single nonlinearity with a gated one (SwiGLU).UnderstandShould know
- PagedAttention and Continuous BatchingPagedAttention stores each request's KV cache in small fixed-size blocks found through a lookup table, like an operating system's virtual memory, so memory is allocated as tokens arrive, wasted space nearly disappears, and requests can share common blocks; with batching at the level of individual decode steps, far more users fit on a GPU.UnderstandShould know
- Pruning and SparsityPruning removes weights (or whole neurons, heads or layers) that contribute little, leaving a sparse network that needs less storage and, if the hardware can exploit the pattern, less computation.UnderstandShould know
- Rotary Position Embedding (RoPE)RoPE encodes a token's position by rotating each pair of coordinates in its query and key by an angle proportional to the position, so the dot product between a query and a key depends only on how far apart the two tokens are.UnderstandShould know
- Speculative DecodingSpeculative decoding lets a small draft model guess several tokens ahead and has the large model check all the guesses in one parallel pass, keeping the ones it agrees with by a rule that leaves the output distribution exactly the large model's own.UnderstandShould know
What do I actually need to remember?
- Prefill processes the prompt in parallel; decode makes one token per pass and is limited by memory bandwidth, not arithmetic.
- Each decode step reads every weight and every cached key and value; batching shares the weight reads across users.
- KV cache per token = 2 × layers × KV heads × head size × bytes; it grows with context length and with users.
- Grouped-query attention shares key/value heads across query heads, shrinking the cache 4–16× in Llama 3.
- Attention's score matrix is n × n; FlashAttention computes it exactly in on-chip tiles without ever storing it.
- Modern blocks: pre-norm, RMSNorm, SwiGLU and rotary position embeddings (RoPE).
- Mixture of experts: a router sends each token to a few experts; memory follows total parameters, compute follows active ones.
- Quantization stores weights in 8 or 4 bits with shared scales; small groups and outlier handling keep the error down.
- LoRA trains a low-rank update on a frozen model; QLoRA freezes the model in 4 bits.
- Speculative decoding drafts tokens cheaply and verifies them in one pass, without changing the output distribution.
You do not need to memorize everything else. This list is the revision sheet.
Key papers
Efficiently Scaling Transformer Inference
Reiner Pope, Sholto Douglas et al. · 2022
The clearest account of why serving a large model is a memory problem: it splits inference into prefill and decode and shows when loading weights or the KV cache dominates.
- Problem
- Generating text one token at a time from a 500B-parameter model under tight latency targets uses the hardware badly, and the right way to split the model across chips depends on the workload.
- What was new
- A simple analytical model of inference cost (memory time vs compute time), partitioning layouts chosen with it, and the observation that multi-query attention's smaller cache allows much longer contexts.
How to read it: Read Section 2 (inference cost and the memory-time analysis). The worked example of a 3 TB KV cache is in Section 2.
Fast Transformer Decoding: One Write-Head is All You Need
Noam Shazeer · 2019
Introduced multi-query attention: all heads share one key and one value, which shrinks the cache that every decoding step must reload.
- Problem
- Incremental decoding is slow because each step reloads the large key and value tensors from memory.
- What was new
- Keep many query heads but a single key head and value head. In the paper's translation model, decoder time per token fell from 46 µs to 3.8 µs with a small quality loss.
- Built on
- Attention Is All You Need
How to read it: Short and readable. Sections 2.4 and 3.1 count memory accesses against arithmetic; Table 2 has the timings.
GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints
Joshua Ainslie, James Lee-Thorp et al. · 2023
Grouped-query attention is the compromise most open LLMs now use: a few key/value heads shared by groups of query heads.
- Problem
- Multi-query attention is fast but can lose quality and train less stably, and retraining a model just for faster inference is expensive.
- What was new
- Groups of query heads share key/value heads (between one and all), and existing multi-head checkpoints can be converted with about 5% of the original pretraining compute.
How to read it: Figure 2 shows the three attention types side by side; Table 1 gives quality against inference time.
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
Tri Dao, Daniel Y. Fu et al. · 2022
Showed that attention was slow because of memory traffic, not arithmetic, and fixed it exactly: same outputs, far fewer reads and writes, memory linear in sequence length.
- Problem
- Standard attention writes the full n × n score matrix to slow GPU memory and reads it back, so long sequences are slow and memory-hungry. Approximate methods cut arithmetic but rarely saved wall-clock time.
- What was new
- Compute attention in tiles that stay in fast on-chip memory, using an online softmax, and recompute in the backward pass instead of storing the scores.
How to read it: Figure 1 (the memory hierarchy and the tiling loop) and Algorithm 1 carry the idea; the IO-complexity proofs can wait.
FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
Tri Dao · 2023
A lesson in how much speed hides in work partitioning: the same exact algorithm, about twice as fast.
- Problem
- FlashAttention reached only 25–40% of an A100's theoretical peak, far below an optimized matrix multiplication.
- What was new
- Fewer non-matrix-multiply operations, parallelism across the sequence as well as across heads, and better division of work inside each thread block: 50–73% of peak on an A100.
How to read it: Read after FlashAttention; Section 3 lists the three changes.
On Layer Normalization in the Transformer Architecture
Ruibin Xiong, Yunchang Yang et al. · 2020 · ICML 2020
Explains why modern Transformers put layer normalization before each sublayer ('pre-LN') rather than after it.
- Problem
- The original post-LN Transformer needed a careful learning-rate warm-up to train stably.
- What was new
- Analysis showing pre-LN keeps gradients well-behaved at initialization, allowing training without warm-up.
Root Mean Square Layer Normalization
Biao Zhang, Rico Sennrich · 2019
RMSNorm, LayerNorm without the mean subtraction, is the normalization in LLaMA-style models.
- Problem
- LayerNorm's computation adds noticeable overhead, and it was unclear whether re-centring each vector was needed at all.
- What was new
- Divide by the root mean square only. Quality was comparable to LayerNorm while running time fell by 7–64% across the models tested.
- Built on
- Layer Normalization
How to read it: Equations 2–4 are the whole method; the rest is experiments.
GLU Variants Improve Transformer
Noam Shazeer · 2020
The origin of SwiGLU, the gated feed-forward layer used by LLaMA, Mistral and many others.
- Problem
- The Transformer's feed-forward layer applies one fixed nonlinearity (ReLU or GELU) between two matrices; could a gate do better?
- What was new
- Multiply one linear projection by an activated second projection (a gated linear unit), shrinking the hidden width so parameter count stays equal. Several variants, including SwiGLU, improved quality.
- Built on
- Attention Is All You Need
How to read it: Four pages. Equations 5–6 define the variants; note the author's own admission that the paper offers results, not an explanation.
RoFormer: Enhanced Transformer with Rotary Position Embedding
Jianlin Su, Yu Lu et al. · 2021
Rotary position embedding (RoPE) is how most open LLMs encode word order: rotate queries and keys by position-dependent angles so attention scores depend on relative distance.
- Problem
- Added position vectors mix position into the content of each token, and attention scores don't directly see how far apart two tokens are.
- What was new
- Encode absolute position as a rotation of each pair of query and key coordinates; the dot product of two rotated vectors then depends only on their offset.
- Built on
- Attention Is All You Need
How to read it: Section 3.2.1 has the 2-D case that makes the idea clear; the general form is the same rotation applied to every pair of dimensions.
Extending Context Window of Large Language Models via Positional Interpolation
Shouyuan Chen, Sherman Wong et al. · 2023
A simple trick for using RoPE models beyond their trained length: squeeze new positions into the old range instead of extrapolating.
- Problem
- RoPE models break down on positions longer than they were trained on.
- What was new
- Linearly scale down position indices to fit the original window, then fine-tune briefly: LLaMA models were extended to 32,768 tokens within 1,000 steps.
Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
Noam Shazeer, Azalia Mirhoseini et al. · 2017
Made conditional computation work at scale: a learned gate sends each input to a few of many expert networks, so parameters can grow far faster than compute.
- Problem
- A network's capacity grows with its parameters, but in a dense network every parameter costs computation on every input.
- What was new
- A trainable gate that picks a sparse top-k set of up to thousands of feed-forward experts, with noise and auxiliary losses to balance their use. MoE layers of up to 137B parameters were applied between LSTM layers.
How to read it: Section 2 (the gating network) and Section 4 (balancing expert use) are the parts still in use today.
Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
William Fedus, Barret Zoph, Noam Shazeer · 2021
Simplified mixture-of-experts Transformers (route each token to one expert) and made them train stably, including in bfloat16.
- Problem
- Mixture-of-experts models were complicated, communication-heavy and unstable to train.
- What was new
- Top-1 routing, an expert capacity limit with a simple load-balancing loss, and stability tricks. Up to 7× faster pretraining than T5-Base for the same compute, and models up to a trillion parameters.
- Influenced
- Mixtral of Experts
How to read it: Section 2 (routing, capacity factor and the balancing loss in Equation 4) is the core; the rest is scaling experiments.
Mixtral of Experts
Albert Q. Jiang, Alexandre Sablayrolles et al. · 2024
A widely used open mixture-of-experts LLM, and a clean example of total versus active parameters: 47B stored, 13B used per token.
- Problem
- Can a sparse model match much larger dense models while computing far less per token?
- What was new
- Mistral 7B's architecture with 8 feed-forward experts per layer and top-2 routing. It matched or beat Llama 2 70B and GPT-3.5 on the reported benchmarks.
How to read it: Section 2 is a page long and covers the whole architecture; Section 5 looks at which experts tokens actually choose.
DeepSeek-V3 Technical Report
DeepSeek-AI et al. · 2024
A detailed report on a frontier-scale open mixture-of-experts model that brings this chapter's ideas together: sparse experts, a compressed attention cache, FP8 training.
- Problem
- How to train and serve a very large model at a fraction of the usual cost.
- What was new
- 671B total parameters with 37B active per token (256 routed experts plus a shared one, 8 chosen per token), multi-head latent attention to shrink the KV cache, load balancing without an auxiliary loss, and FP8 mixed-precision training.
How to read it: Read Section 2 (architecture) first; Section 3 on infrastructure is valuable if you work on systems.
LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale
Tim Dettmers, Mike Lewis et al. · 2022
Found why naive 8-bit quantization breaks large models (a few huge outlier features) and how to work around them, halving inference memory without degradation.
- Problem
- Quantizing models above a few billion parameters to 8 bits ruined accuracy, for reasons that were not understood.
- What was new
- Vector-wise quantization with separate scales, plus a mixed-precision decomposition that keeps the rare outlier dimensions in 16 bits while over 99.9% of values are multiplied in 8 bits. Outliers appeared in all layers from about 6.7B parameters.
- Built on
- Mixed Precision Training
How to read it: Figure 1 (accuracy collapse at 6.7B) and Section 4 on emergent outlier features are the memorable parts.
GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
Elias Frantar, Saleh Ashkboos et al. · 2022
Made 3–4-bit weights practical for very large models without retraining, which is how many open models are run on single GPUs.
- Problem
- Rounding each weight to the nearest low-bit value independently loses too much accuracy at 3–4 bits.
- What was new
- Quantize weights one column at a time and adjust the not-yet-quantized weights to compensate, using approximate second-order information from a little calibration data. A 175B model took about four GPU hours.
How to read it: Section 4 builds the algorithm step by step; Figure 2 and Algorithm 1 show the block-by-block procedure.
AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration
Ji Lin, Jiaming Tang et al. · 2023
Shows that which weights matter depends on the activations they meet, and protects those with a simple rescaling.
- Problem
- Low-bit weight quantization hurts some weights much more than others; keeping important ones in higher precision is awkward for hardware.
- What was new
- Find the roughly 1% of salient weight channels from activation statistics and scale them up before quantizing, so they lose less precision, with no backpropagation.
How to read it: Section 3.1 (protecting 1% of weights) and 3.2 (the scaling trick) are the method.
FP8 Formats for Deep Learning
Paulius Micikevicius, Dusan Stosic et al. · 2022
Specified the two 8-bit floating-point formats (E4M3 and E5M2) that recent GPUs support for training and inference.
- Problem
- 16-bit training was standard; could 8-bit floating point keep quality while doubling throughput?
- What was new
- Two encodings, one with more precision (4 exponent bits, 3 mantissa bits) and one with more range (5 and 2), matching 16-bit training quality on models up to 175B parameters.
- Built on
- Mixed Precision Training
- Influenced
- DeepSeek-V3 Technical Report
Distilling the Knowledge in a Neural Network
Geoffrey Hinton, Oriol Vinyals, Jeff Dean · 2015
Defined knowledge distillation: train a small student on a large teacher's softened output probabilities, which carry more information than hard labels.
- Problem
- Large models and ensembles predict well but are too expensive to deploy widely.
- What was new
- Raise the softmax temperature so the teacher's probabilities for wrong answers become visible, and train the student to match them at the same temperature.
How to read it: Section 2 explains temperature and soft targets in a page.
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut et al. · 2019
A widely used demonstration that distillation during pretraining produces a much cheaper general-purpose language model.
- Problem
- BERT-sized models were too heavy for many deployments.
- What was new
- Distill during pretraining with a combined loss: 40% smaller, 60% faster, retaining 97% of BERT's language-understanding score.
Learning both Weights and Connections for Efficient Neural Networks
Song Han, Jeff Pool et al. · 2015
The classic train–prune–retrain recipe that showed most connections in a trained network can be removed.
- Problem
- Networks are larger than they need to be, which costs memory, energy and time on devices.
- What was new
- Train, remove small-magnitude connections, then retrain the rest. AlexNet went from 61 million to 6.7 million parameters, and VGG-16 from 138 to 10.3 million, without losing accuracy.
The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks
Jonathan Frankle, Michael Carbin · 2018
A provocative finding about why pruning works: dense networks contain small subnetworks that could have been trained alone.
- Problem
- Pruned networks perform well, but training the same sparse architecture from scratch usually doesn't.
- What was new
- Reset the surviving weights to their original initial values and retrain: winning tickets under 10–20% of the original size matched the full network on MNIST and CIFAR-10.
How to read it: Read the introduction and Section 2. The experiments are on small vision networks; whether it carries over to LLMs is a separate question.
SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot
Elias Frantar, Dan Alistarh · 2023
Showed that GPT-scale models can lose half their weights in one pass, without retraining.
- Problem
- Pruning methods that need retraining are impractical for 100B+ parameter models.
- What was new
- A one-shot pruning solver that reached 50–60% unstructured sparsity on OPT-175B and BLOOM-176B in under 4.5 hours, with a small perplexity increase.
Parameter-Efficient Transfer Learning for NLP
Neil Houlsby, Andrei Giurgiu et al. · 2019
Introduced adapters, small trainable layers inserted into a frozen network: the start of parameter-efficient fine-tuning.
- Problem
- Fine-tuning a separate copy of a large model for every task is wasteful when tasks arrive one after another.
- What was new
- Insert small bottleneck layers into each Transformer block and train only those. On GLUE, within 0.4% of full fine-tuning while adding 3.6% parameters per task.
How to read it: Figure 2 shows where the adapters go; that picture is most of the paper.
LoRA: Low-Rank Adaptation of Large Language Models
Edward J. Hu, Yelong Shen et al. · 2021
The standard way to adapt a large model cheaply: freeze its weights and learn a small low-rank update for some matrices.
- Problem
- Full fine-tuning updates and stores every parameter for every task: for GPT-3 175B, a 350 GB checkpoint per task.
- What was new
- Represent each weight update as the product of two thin matrices (rank r), train only those, and merge them into the weights afterwards, so there is no extra inference latency. For GPT-3 it cut trainable parameters about 10,000-fold and GPU memory about threefold.
How to read it: Section 4 (the method, one page) and Section 7 (why a low rank is enough) are the essentials.
QLoRA: Efficient Finetuning of Quantized LLMs
Tim Dettmers, Artidoro Pagnoni et al. · 2023
Combined 4-bit quantization with LoRA so a 65B model can be fine-tuned on one 48 GB GPU, putting large-model fine-tuning within reach of small labs.
- Problem
- Even with LoRA, a large model's 16-bit frozen weights need several GPUs just to hold them during fine-tuning.
- What was new
- Freeze the base model in a 4-bit NormalFloat format, backpropagate through it into LoRA adapters, quantize the quantization constants too (double quantization), and page optimizer states to CPU memory during spikes.
How to read it: Section 3 explains NF4, double quantization and paged optimizers in two pages; the block-size arithmetic (0.5 → 0.127 bits per parameter) is worth checking by hand.
Fast Inference from Transformers via Speculative Decoding
Yaniv Leviathan, Matan Kalman, Yossi Matias · 2022
Speeds up generation without changing its output distribution: a small model guesses ahead and the large model checks the guesses in parallel.
- Problem
- Generating K tokens takes K sequential runs of the large model, and each run leaves most of the hardware idle.
- What was new
- Draft γ tokens with a cheap model, score them with one parallel pass of the target model, and accept or resample with a rule that provably preserves the target's distribution. 2–3× faster on T5-XXL with identical outputs.
How to read it: Section 2.3 (the accept/resample rule) and Equation 1 (expected tokens per pass) are what the lab in this chapter implements.
Accelerating Large Language Model Decoding with Speculative Sampling
Charlie Chen, Sebastian Borgeaud et al. · 2023
Independently arrived at the same idea and showed it on a 70B model in a distributed setting.
- Problem
- Sampling from a large Transformer is memory-bandwidth bound and slow.
- What was new
- Score short drafts from a faster model in parallel with a modified rejection sampling scheme that preserves the target distribution: a 2–2.5× decoding speed-up on Chinchilla 70B.
Efficient Memory Management for Large Language Model Serving with PagedAttention
Woosuk Kwon, Zhuohan Li et al. · 2023
Brought operating-system paging to the KV cache (the vLLM server), so many more requests fit in GPU memory at once.
- Problem
- Serving systems reserved one contiguous KV-cache region per request for its maximum length; only 20–38% of the reserved memory held actual token states.
- What was new
- Store each request's cache in fixed-size blocks located through a block table, allocated on demand and shareable between requests. Throughput rose 2–4× at the same latency.
How to read it: Section 3 (why existing systems waste memory) and Figure 6 (the block table) are the core.
Llama 2: Open Foundation and Fine-Tuned Chat Models
Hugo Touvron, Louis Martin et al. · 2023
A widely used open model family; its larger models adopted grouped-query attention for faster inference.
- Problem
- Open models lagged behind closed chat models, and serving the larger ones was expensive.
- What was new
- Pretrained models from 7B to 70B with doubled context length and grouped-query attention in the larger models, plus chat versions trained with SFT and RLHF.
- Built on
- LLaMA: Open and Efficient Foundation Language Models; GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints
- Influenced
- The Llama 3 Herd of Models
How to read it: For this chapter, read Section 2.2 and Appendix A.2.1, which compares multi-head, multi-query and grouped-query attention.
The Llama 3 Herd of Models
Aaron Grattafiori, Abhimanyu Dubey et al. · 2024
The most complete public account of building a frontier-scale model end to end: data pipeline, scaling-law experiments, 16,384-GPU training, failures and all.
- Problem
- Frontier labs had largely stopped publishing how their models were built.
- What was new
- A 405B model trained on 15.6T tokens with 3.8 × 10²⁵ FLOPs, with detailed data cleaning, a data mix chosen by scaling experiments, 4D parallelism at 38–43% utilisation, and 466 job interruptions in a 54-day period.
How to read it: It is 90+ pages. For this chapter read Section 3 (pre-training) only: data, scaling laws, infrastructure and the training recipe.
Watch
Stanford Online
Stanford CS336 Language Modeling from Scratch | Spring 2025 | Lecture 10: Inference
The closest single lecture to this chapter: the arithmetic of serving a language model and the main ways to make it cheaper.
Covers: Prefill and generation, arithmetic intensity, the KV cache and how GQA shrinks it, quantization, pruning and distillation, speculative sampling, continuous batching and paged attention.
Stanford Online
Stanford CS336 Lang. Modeling from Scratch | Spring 2025 | Lec. 3: Architectures, Hyperparameters
A survey of what changed inside the Transformer block between 2017 and today's open models, and which choices nearly everyone now agrees on.
Covers: Pre-norm versus post-norm, RMSNorm, gated activations such as SwiGLU, rotary position embeddings, multi-query and grouped-query attention, and common hyperparameter choices.
Stanford Online
Stanford CS336 Language Modeling from Scratch | Spring 2025 | Lecture 4: Mixture of experts
A careful lecture on why mixture-of-experts models are now common at the frontier, and on what makes them hard to train.
Covers: Routers and top-k routing, expert and device load balancing, token dropping, training stability, and recent designs such as DeepSeek's.
Stanford Online
Stanford CS336 I Language Modeling from Scratch | Spring 2025 | Lecture 5: GPUs
The hardware background for this chapter: why memory movement, not arithmetic, so often sets the speed.
Covers: How GPUs execute work, the memory hierarchy, arithmetic intensity and the roofline model, tiling, and how FlashAttention applies these ideas.
What came next?
Chapter 12
Embeddings, RAG & the LLM Application Stack
A model's knowledge is frozen at training time and can't cite its sources.
This chapter is being written.