Skip to content
Road to Intelligence

Part III · Large Language Models

Chapter 11

Inside Modern LLMs

The engineering that makes large models fast, cheap and adaptable.

2 h 30 min core path14 concepts3 interactivesCore path · 7 optional concepts foldedDeep · all 14 concepts shown in full

In one sentenceKV caches, efficient attention, mixture-of-experts, quantization and low-rank adaptation each trade something away to fit large models onto real hardware.

The problem

The serving bill

Chapter 9 trained Llama 3 405B. Chapter 10 turned a checkpoint like it into an assistant. Now someone has to run it, for many people at once, quickly and at a price they'll pay. Training happens once; inference happens every time anyone asks anything.

Meta's Llama 3 report describes this part too, and it makes a good running example:

WhatLlama 3 405BSection
Weights in 16 bits810 GB: more than one server's eight 80 GB GPUsThe serving bill
Cache per token of contextabout 0.5 MB, thanks to grouped-query attentionKV cache; Fewer keys and values
Longest context128K tokens, about 68 GB of cache for one such conversationThe quadratic wall
Served on16 GPUs across two servers in BF16The serving bill
FP8 weights and activations in the feed-forward layersup to 50% more prefill throughputFewer bits per weight

The report states that in BF16 the 405B model does not fit in the memory of a single machine with 8 H100 GPUs, so Meta ran inference across 16 GPUs on two machines, and that FP8 quantization improved prefill throughput by up to 50% Established. The 810 GB, 0.5 MB and 68 GB are our arithmetic from the published architecture, explained below.

Inference has two phases. Prefill reads the prompt in one parallel pass, much like training. Decode then produces one token per pass, and every pass must read all the weights from memory. An H100 reads memory at 3.35 TB/s Established and can do about 295 dense BF16 operations in the time it reads one byte. A decode step for one user does about one operation per byte of weights. So during decoding the GPU's arithmetic units mostly wait: the speed limit is memory bandwidth, not FLOPs.

That observation organizes the whole chapter. Each idea below reduces the bytes moved per generated token, or the memory that limits how many users share a pass.

The first saving

Remember, don't recompute

To generate the next token, attention needs keys and values for every token before it. Those don't change as the text grows, so the model computes them once and stores them: the KV cache. Each new token then needs only its own query, key and value.

The cache isn't small. Per token it holds a key and a value for every layer and every key/value head. For OPT-13B that is 800 KB per token, so one 2,048-token request needs up to 1.6 GB Established. Unlike the weights, each conversation has its own cache, and every decode step reads all of it.

Try it below with real model configurations. With Llama 3 8B and one user at 8,192 tokens, the 16 GB of weights dominate, and memory bandwidth caps decoding at about 196 tokens per second. Raise the users to 32: each gets about 67 tokens per second, but the GPU produces over 2,100 in total, because the weights are read once per step for everyone. Now switch the attention to "one per query head": the cache grows to 137 GB and nothing fits.

Try it

How Big Is the Cache?

Serve real Llama models to many users at once: add up weights and KV cache against GPU memory, change the number of key/value heads, and see the memory-bandwidth limit on tokens per second.

Know well8 min

The second saving

Fewer keys and values

The cache grows with the number of key/value heads, and the fix is to have fewer of them. Multi-query attention keeps every query head but shares one key head and one value head among them; in its 2019 translation experiments, decoder time per token fell from 46 µs to 3.8 µs with only minor quality loss Established. Grouped-query attention sits in between: groups of query heads share a key/value head. Converted from multi-head checkpoints with about 5% of the original pretraining compute, GQA with 8 groups came close to multi-head quality at close to multi-query speed Established.

Llama 3 uses 8 key/value heads at every size Established. The result is striking: the 405B model stores about 504 KiB of cache per token, slightly less than Llama 2 7B's 512 KiB with full multi-head attention, despite having 58 times as many parameters.

Long contexts

The quadratic wall

Chapter 7 noted that attention compares every token with every other: nn tokens give an n×nn \times n score matrix per head. At 8,192 tokens that is 67 million scores per head per layer. Double the context and the scores quadruple.

For years the standard implementation wrote that whole matrix to the GPU's main memory, read it back for the softmax, and read it again to multiply by the values. FlashAttention never writes it. It processes attention in tiles small enough for the GPU's fast on-chip memory, keeping a running maximum and sum so the softmax can be completed tile by tile. The output is exactly the same. It trained GPT-2 3× faster at 1K tokens and made attention's memory linear in sequence length Established. The arithmetic was never the bottleneck; the trips to memory were.

OptionalFlashAttention· folded on the core path. Open it, or switch to Deep to show it here.

Small changes that stuck

The modern block

Open the code of a recent open model and you'll recognize Chapter 7's block, with a few edits. Llama 2 describes its architecture as the standard Transformer with pre-normalization using RMSNorm, the SwiGLU activation and rotary positional embeddings Established:

  • Pre-norm: normalize at the start of each sublayer, not after adding the residual, so the residual path is a clean sum and training is less fragile.
  • RMSNorm: divide by the root mean square and skip LayerNorm's mean subtraction; cheaper, and in practice just as good.
  • SwiGLU: a gated feed-forward layer, one projection multiplied by an activated second one. Its author found it improved quality and offered no explanation for why Established.
OptionalThe Modern Decoder Block: Pre-Norm, RMSNorm, SwiGLU· folded on the core path. Open it, or switch to Deep to show it here.

The fourth edit is position. Instead of adding a position vector to each token, RoPE rotates queries and keys by angles proportional to their positions, so attention scores depend on how far apart two tokens are. Its frequencies are also the main dial for stretching a model's context beyond what it was trained on.

OptionalRotary Position Embedding (RoPE)· folded on the core path. Open it, or switch to Deep to show it here.

Parameters without the compute

Many experts, few at a time

A dense model uses every parameter for every token. A mixture-of-experts layer replaces the feed-forward network with many smaller experts and a router that sends each token to a few of them. Memory holds all the experts; each token pays for only the ones it uses.

Mixtral 8x7B routes each token to 2 of 8 experts per layer: 47B parameters in total, 13B used per token. It matched or beat Llama 2 70B and GPT-3.5 on the benchmarks its authors report Established. DeepSeek-V3 has 671B parameters, of which 37B are activated for each token Established. Step through one token's trip through a layer:

One mixture-of-experts layer · 8 experts, top-2 routing
Token:
  1. Expert 11.2idle
  2. Expert 2-0.3idle
  3. Expert 30.52runs
  4. Expert 40.1idle
  5. Expert 50.4idle
  6. Expert 6-1.0idle
  7. Expert 70.48runs
  8. Expert 80.0idle

Output = 0.52 × expert 3 + 0.48 × expert 7. Six experts do no work for this token, but all eight stay in memory. Try another token: it can choose different experts.

Mixtral 8x7B, per token

stored (memory)47B
used (compute)13B

Router scores are made up for illustration; the top-2 selection and softmax are computed. Mixtral’s parameter counts are from its paper.

The savings are in arithmetic, not memory: Mixtral's authors note that its serving memory is proportional to its 47B total parameters Established. MoE also brings its own engineering: keeping the experts evenly used, and moving tokens between the GPUs that hold different experts.

Smaller numbers

Fewer bits per weight

If decoding is limited by bytes, store fewer bytes per weight. Quantization replaces each 16-bit weight by an 8- or 4-bit integer and a scale shared by a group of weights. Llama 3 70B's weights take 140 GB in 16 bits and about 39 GB at 4 bits, small enough for a single 80 GB GPU.

The difficulty is a few very large values. A shared scale has to stretch to cover the largest number in its group, and everything else gets squeezed into a few levels. Dettmers and colleagues found that from about 6.7B parameters, large language models develop systematic outlier features that ruin naive 8-bit quantization Established, and handled them by keeping those few dimensions in 16 bits.

See it on a toy row of 512 weights below. In 8 bits with one scale for the row, the error is under 1%. Drop to 4 bits: about 11%. Add one outlier: at 4 bits, 99.6% of the weights now round to zero. Switch to one scale per 64 weights and the damage is confined to the outlier's own group, at a cost of half a bit per weight for the extra scales.

Try it · toy model

Squeeze the Weights

Round a row of weights to 8, 4, 3 or 2 bits, add one outlier, and see why a shared scale fails and per-group scales rescue it.

Know well7 min

Better methods round more cleverly. GPTQ quantized 175B-parameter models to 3–4 bits in about four GPU hours by adjusting the remaining weights to compensate for each rounding error Established, and AWQ protects the roughly 1% of weights that matter most for the activations they meet.

Compression

Smaller models from bigger ones

Two older ideas make models smaller in different ways. Distillation trains a small student to match a large teacher's output probabilities, which say more than the right answer alone: that a 2 looks a bit like a 7. DistilBERT was 40% smaller and 60% faster than BERT while keeping 97% of its language-understanding performance Established.

OptionalKnowledge Distillation· folded on the core path. Open it, or switch to Deep to show it here.

Pruning removes weights that contribute little. Trained networks tolerate a lot of it, but zeros only save time when the hardware can skip them, so patterns matter more than counts.

OptionalPruning and Sparsity· folded on the core path. Open it, or switch to Deep to show it here.

Adapting cheaply

Adapting without retraining

Fine-tuning every weight of a large model, as in Chapter 10, needs memory for gradients and optimizer states on every parameter, and produces a full copy of the model per task. LoRA freezes the model and learns, for selected matrices, a low-rank update: two thin matrices whose product is added to the weights. Slide the rank and watch the number of trained parameters:

LoRA: a frozen matrix plus a low-rank update
W0frozen · 4,096 × 4,096+B×A · 8 × 4,096trainable · rank 8
This matrix, full fine-tuning
16,777,216
This matrix, LoRA
65,536 (0.39%)
Llama 3 8B, query + value in all 32 layers
3.4M (0.043% of 8B)

Trainable numbers are r × (d + k) per matrix. B starts at zero, so training begins from the unchanged model. Llama 3 8B’s value projection is 4,096 × 1,024 because of grouped-query attention. Counts are our arithmetic from the published architecture; Adam’s states are needed only for the trainable part.

On GPT-3 175B, LoRA cut trainable parameters 10,000-fold and GPU memory 3-fold compared with full fine-tuning with Adam, while matching its quality, and the update can be merged into the weights so inference is no slower Established. QLoRA keeps the frozen model in 4 bits and fine-tuned a 65B-parameter model on a single 48 GB GPU Established.

Using the idle arithmetic

Guess ahead, check in parallel

Decoding leaves arithmetic idle, and checking five tokens costs about as much as generating one. Speculative decoding puts that to use: a small draft model guesses several tokens ahead, and the large model scores all of them in one pass, keeping each guess by a rule that, remarkably, leaves the output distribution exactly the large model's own.

In the lab, the target model prefers "the" (0.5) and the draft prefers "a" (0.4). Run thousands of draft-and-verify steps: the words come out at the target's proportions, not the draft's, and the draft was accepted 80% of the time. Then set the acceptance rate, draft length and draft cost to see the expected speed-up.

Try it · toy model

Draft and Verify

Run speculative decoding's accept-or-resample rule thousands of times and watch the output keep the large model's distribution, then trade acceptance rate, draft length and draft cost for speed.

Understand8 min

Speculative decoding ran T5-XXL 2–3× faster with identical outputs Established.

OptionalSpeculative Decoding· folded on the core path. Open it, or switch to Deep to show it here.

The server

Serving many users

Batching is the most powerful lever of all, because each pass's weight reads are shared by every request in the batch. What limits the batch is cache memory. The vLLM authors measured that existing servers used only 20.4–38.2% of their KV-cache memory for actual tokens Established; the rest was reserved for replies that might grow, or lost to fragmentation.

PagedAttention borrows virtual memory from operating systems: the cache lives in small blocks allocated as tokens arrive and found through a table, and requests with a common prompt share blocks. Together with batching at the level of individual decode steps, so finished requests leave and new ones join at every step, the vLLM server improved throughput 2–4× at the same latency over the systems it was compared with Established.

Making models cheaper to run, 2015–2024Open the full timeline →
The deep-learning revolutionTransformers and scaleLLMs, agents and reasoning (evolving)2015: Deep Q-networks master AtariDeep Q-networks master Atari20152015: ResNetResNet20152016: AlphaGo beats Lee SedolAlphaGo beats Lee Sedol20162017: Sparsely-gated mixture of expertsSparsely-gated mixture of experts20172017: The TransformerThe Transformer20172017: Learning from human preferencesLearning from human preferences20172018: GPT-1: generative pretrainingGPT-1: generative pretraining20182018: BERTBERT20182019: GPT-2GPT-220192019: ZeRO removes redundant copiesZeRO removes redundant copies20192020: Scaling lawsScaling laws20202020: GPT-3 and in-context learningGPT-3 and in-context learning20202020: Summaries learned from human feedbackSummaries learned from human feedback20202020: Vision TransformerVision Transformer20202020: AlphaFold 2AlphaFold 220202021: CLIP: images meet languageCLIP: images meet language20212021: LoRA: low-rank fine-tuningLoRA: low-rank…20212022: Chain-of-thought promptingChain-of-thought…20222022: InstructGPT and RLHFInstructGPT and RLHF20222022: Open text-to-image modelsOpen text-to…20222022: ChatGPTChatGPT20222022: Constitutional AIConstitutional AI20222023: GPT-4GPT-420232023: Direct Preference OptimizationDirect Preference Optimization20232023: Capable open-weight LLMsCapable open-weight LLMs20232023: PagedAttention and vLLMPagedAttentio…20232024: Mixtral: an open mixture-of-experts modelMixtral: an open…20242024: The Llama 3 report opens up a frontier-scale runThe Llama 3 report…20242024: EU AI Act enters into forceEU AI Act enters…20242024: Reasoning models: OpenAI o1Reasoning models: OpenAI o12024

The bigger picture

Why it matters

None of this changes what a language model is. The model of Chapter 8 still predicts the next token; the block of Chapter 7 is still recognizably there. What changed is the cost of each token, and cost decides what gets built: longer contexts, many samples per question, models on laptops and phones, fine-tuned variants for every customer.

For an engineer, the habit to take away is the napkin calculation: bytes per token, bytes per second, and which of weights, cache or communication dominates. It predicts which technique will help before you try it. For a researcher, the lesson is that architecture now answers to hardware. Grouped-query attention, mixture-of-experts and FlashAttention are judged as much by memory traffic as by loss. Inference cost is increasingly a design constraint at training time too, as Chapter 9's "train smaller, longer" argument showed Interpretation.

Each technique trades something: quantization and distillation give up a little quality; mixture-of-experts gives up memory and simplicity; speculative decoding spends arithmetic; LoRA restricts what an update can express. None is free, and published gains are measured on particular models and tasks.

What comes next

What comes next

A served model is fast and cheap, but its knowledge is still frozen at the end of training, and it can't say where an answer came from. Chapter 12 connects a model to documents it has never seen: embeddings that measure meaning, retrieval that finds the relevant passages, and the application stack around them. The KV cache returns there too, as the cost of every retrieved passage you put in the context.

Concepts in this chapter

Mark each one as you go. Must-know concepts are the core path.

What do I actually need to remember?

  • Prefill processes the prompt in parallel; decode makes one token per pass and is limited by memory bandwidth, not arithmetic.
  • Each decode step reads every weight and every cached key and value; batching shares the weight reads across users.
  • KV cache per token = 2 × layers × KV heads × head size × bytes; it grows with context length and with users.
  • Grouped-query attention shares key/value heads across query heads, shrinking the cache 4–16× in Llama 3.
  • Attention's score matrix is n × n; FlashAttention computes it exactly in on-chip tiles without ever storing it.
  • Modern blocks: pre-norm, RMSNorm, SwiGLU and rotary position embeddings (RoPE).
  • Mixture of experts: a router sends each token to a few experts; memory follows total parameters, compute follows active ones.
  • Quantization stores weights in 8 or 4 bits with shared scales; small groups and outlier handling keep the error down.
  • LoRA trains a low-rank update on a frozen model; QLoRA freezes the model in 4 bits.
  • Speculative decoding drafts tokens cheaply and verifies them in one pass, without changing the output distribution.

You do not need to memorize everything else. This list is the revision sheet.

Key papers

Essential

Efficiently Scaling Transformer Inference

Reiner Pope, Sholto Douglas et al. · 2022

The clearest account of why serving a large model is a memory problem: it splits inference into prefill and decode and shows when loading weights or the KV cache dominates.

Problem
Generating text one token at a time from a 500B-parameter model under tight latency targets uses the hardware badly, and the right way to split the model across chips depends on the workload.
What was new
A simple analytical model of inference cost (memory time vs compute time), partitioning layouts chosen with it, and the observation that multi-query attention's smaller cache allows much longer contexts.

How to read it: Read Section 2 (inference cost and the memory-time analysis). The worked example of a 3 TB KV cache is in Section 2.

~1 h readarXiv:2211.05102✓ verified 2026-10-05
Important

Fast Transformer Decoding: One Write-Head is All You Need

Noam Shazeer · 2019

Introduced multi-query attention: all heads share one key and one value, which shrinks the cache that every decoding step must reload.

Problem
Incremental decoding is slow because each step reloads the large key and value tensors from memory.
What was new
Keep many query heads but a single key head and value head. In the paper's translation model, decoder time per token fell from 46 µs to 3.8 µs with a small quality loss.

How to read it: Short and readable. Sections 2.4 and 3.1 count memory accesses against arithmetic; Table 2 has the timings.

~25 min readarXiv:1911.02150✓ verified 2026-10-05
Essential

GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Joshua Ainslie, James Lee-Thorp et al. · 2023

Grouped-query attention is the compromise most open LLMs now use: a few key/value heads shared by groups of query heads.

Problem
Multi-query attention is fast but can lose quality and train less stably, and retraining a model just for faster inference is expensive.
What was new
Groups of query heads share key/value heads (between one and all), and existing multi-head checkpoints can be converted with about 5% of the original pretraining compute.

How to read it: Figure 2 shows the three attention types side by side; Table 1 gives quality against inference time.

~25 min readarXiv:2305.13245✓ verified 2026-10-05
Essential

FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness

Tri Dao, Daniel Y. Fu et al. · 2022

Showed that attention was slow because of memory traffic, not arithmetic, and fixed it exactly: same outputs, far fewer reads and writes, memory linear in sequence length.

Problem
Standard attention writes the full n × n score matrix to slow GPU memory and reads it back, so long sequences are slow and memory-hungry. Approximate methods cut arithmetic but rarely saved wall-clock time.
What was new
Compute attention in tiles that stay in fast on-chip memory, using an online softmax, and recompute in the backward pass instead of storing the scores.

How to read it: Figure 1 (the memory hierarchy and the tiling loop) and Algorithm 1 carry the idea; the IO-complexity proofs can wait.

~50 min readarXiv:2205.14135✓ verified 2026-10-05
Optional

FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Tri Dao · 2023

A lesson in how much speed hides in work partitioning: the same exact algorithm, about twice as fast.

Problem
FlashAttention reached only 25–40% of an A100's theoretical peak, far below an optimized matrix multiplication.
What was new
Fewer non-matrix-multiply operations, parallelism across the sequence as well as across heads, and better division of work inside each thread block: 50–73% of peak on an A100.

How to read it: Read after FlashAttention; Section 3 lists the three changes.

~30 min readarXiv:2307.08691✓ verified 2026-10-05
Optional

On Layer Normalization in the Transformer Architecture

Ruibin Xiong, Yunchang Yang et al. · 2020 · ICML 2020

Explains why modern Transformers put layer normalization before each sublayer ('pre-LN') rather than after it.

Problem
The original post-LN Transformer needed a careful learning-rate warm-up to train stably.
What was new
Analysis showing pre-LN keeps gradients well-behaved at initialization, allowing training without warm-up.
~1 h readarXiv:2002.04745✓ verified 2026-09-26
Optional

Root Mean Square Layer Normalization

Biao Zhang, Rico Sennrich · 2019

RMSNorm, LayerNorm without the mean subtraction, is the normalization in LLaMA-style models.

Problem
LayerNorm's computation adds noticeable overhead, and it was unclear whether re-centring each vector was needed at all.
What was new
Divide by the root mean square only. Quality was comparable to LayerNorm while running time fell by 7–64% across the models tested.

How to read it: Equations 2–4 are the whole method; the rest is experiments.

~20 min readarXiv:1910.07467✓ verified 2026-10-05
Optional

GLU Variants Improve Transformer

Noam Shazeer · 2020

The origin of SwiGLU, the gated feed-forward layer used by LLaMA, Mistral and many others.

Problem
The Transformer's feed-forward layer applies one fixed nonlinearity (ReLU or GELU) between two matrices; could a gate do better?
What was new
Multiply one linear projection by an activated second projection (a gated linear unit), shrinking the hidden width so parameter count stays equal. Several variants, including SwiGLU, improved quality.

How to read it: Four pages. Equations 5–6 define the variants; note the author's own admission that the paper offers results, not an explanation.

~10 min readarXiv:2002.05202✓ verified 2026-10-05
Important

RoFormer: Enhanced Transformer with Rotary Position Embedding

Jianlin Su, Yu Lu et al. · 2021

Rotary position embedding (RoPE) is how most open LLMs encode word order: rotate queries and keys by position-dependent angles so attention scores depend on relative distance.

Problem
Added position vectors mix position into the content of each token, and attention scores don't directly see how far apart two tokens are.
What was new
Encode absolute position as a rotation of each pair of query and key coordinates; the dot product of two rotated vectors then depends only on their offset.

How to read it: Section 3.2.1 has the 2-D case that makes the idea clear; the general form is the same rotation applied to every pair of dimensions.

~35 min readarXiv:2104.09864✓ verified 2026-10-05
Optional

Extending Context Window of Large Language Models via Positional Interpolation

Shouyuan Chen, Sherman Wong et al. · 2023

A simple trick for using RoPE models beyond their trained length: squeeze new positions into the old range instead of extrapolating.

Problem
RoPE models break down on positions longer than they were trained on.
What was new
Linearly scale down position indices to fit the original window, then fine-tune briefly: LLaMA models were extended to 32,768 tokens within 1,000 steps.
~25 min readarXiv:2306.15595✓ verified 2026-10-05
Essential

Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Noam Shazeer, Azalia Mirhoseini et al. · 2017

Made conditional computation work at scale: a learned gate sends each input to a few of many expert networks, so parameters can grow far faster than compute.

Problem
A network's capacity grows with its parameters, but in a dense network every parameter costs computation on every input.
What was new
A trainable gate that picks a sparse top-k set of up to thousands of feed-forward experts, with noise and auxiliary losses to balance their use. MoE layers of up to 137B parameters were applied between LSTM layers.

How to read it: Section 2 (the gating network) and Section 4 (balancing expert use) are the parts still in use today.

~40 min readarXiv:1701.06538✓ verified 2026-10-05
Important

Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity

William Fedus, Barret Zoph, Noam Shazeer · 2021

Simplified mixture-of-experts Transformers (route each token to one expert) and made them train stably, including in bfloat16.

Problem
Mixture-of-experts models were complicated, communication-heavy and unstable to train.
What was new
Top-1 routing, an expert capacity limit with a simple load-balancing loss, and stability tricks. Up to 7× faster pretraining than T5-Base for the same compute, and models up to a trillion parameters.

How to read it: Section 2 (routing, capacity factor and the balancing loss in Equation 4) is the core; the rest is scaling experiments.

~1 h readarXiv:2101.03961✓ verified 2026-10-05
Important

Mixtral of Experts

Albert Q. Jiang, Alexandre Sablayrolles et al. · 2024

A widely used open mixture-of-experts LLM, and a clean example of total versus active parameters: 47B stored, 13B used per token.

Problem
Can a sparse model match much larger dense models while computing far less per token?
What was new
Mistral 7B's architecture with 8 feed-forward experts per layer and top-2 routing. It matched or beat Llama 2 70B and GPT-3.5 on the reported benchmarks.

How to read it: Section 2 is a page long and covers the whole architecture; Section 5 looks at which experts tokens actually choose.

~20 min readarXiv:2401.04088✓ verified 2026-10-05
Important

DeepSeek-V3 Technical Report

DeepSeek-AI et al. · 2024

A detailed report on a frontier-scale open mixture-of-experts model that brings this chapter's ideas together: sparse experts, a compressed attention cache, FP8 training.

Problem
How to train and serve a very large model at a fraction of the usual cost.
What was new
671B total parameters with 37B active per token (256 routed experts plus a shared one, 8 chosen per token), multi-head latent attention to shrink the KV cache, load balancing without an auxiliary loss, and FP8 mixed-precision training.

How to read it: Read Section 2 (architecture) first; Section 3 on infrastructure is valuable if you work on systems.

~1 h 30 min readarXiv:2412.19437✓ verified 2026-10-05
Essential

LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale

Tim Dettmers, Mike Lewis et al. · 2022

Found why naive 8-bit quantization breaks large models (a few huge outlier features) and how to work around them, halving inference memory without degradation.

Problem
Quantizing models above a few billion parameters to 8 bits ruined accuracy, for reasons that were not understood.
What was new
Vector-wise quantization with separate scales, plus a mixed-precision decomposition that keeps the rare outlier dimensions in 16 bits while over 99.9% of values are multiplied in 8 bits. Outliers appeared in all layers from about 6.7B parameters.

How to read it: Figure 1 (accuracy collapse at 6.7B) and Section 4 on emergent outlier features are the memorable parts.

~50 min readarXiv:2208.07339✓ verified 2026-10-05
Important

GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Elias Frantar, Saleh Ashkboos et al. · 2022

Made 3–4-bit weights practical for very large models without retraining, which is how many open models are run on single GPUs.

Problem
Rounding each weight to the nearest low-bit value independently loses too much accuracy at 3–4 bits.
What was new
Quantize weights one column at a time and adjust the not-yet-quantized weights to compensate, using approximate second-order information from a little calibration data. A 175B model took about four GPU hours.

How to read it: Section 4 builds the algorithm step by step; Figure 2 and Algorithm 1 show the block-by-block procedure.

~45 min readarXiv:2210.17323✓ verified 2026-10-05
Optional

AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration

Ji Lin, Jiaming Tang et al. · 2023

Shows that which weights matter depends on the activations they meet, and protects those with a simple rescaling.

Problem
Low-bit weight quantization hurts some weights much more than others; keeping important ones in higher precision is awkward for hardware.
What was new
Find the roughly 1% of salient weight channels from activation statistics and scale them up before quantizing, so they lose less precision, with no backpropagation.

How to read it: Section 3.1 (protecting 1% of weights) and 3.2 (the scaling trick) are the method.

~35 min readarXiv:2306.00978✓ verified 2026-10-05
Optional

FP8 Formats for Deep Learning

Paulius Micikevicius, Dusan Stosic et al. · 2022

Specified the two 8-bit floating-point formats (E4M3 and E5M2) that recent GPUs support for training and inference.

Problem
16-bit training was standard; could 8-bit floating point keep quality while doubling throughput?
What was new
Two encodings, one with more precision (4 exponent bits, 3 mantissa bits) and one with more range (5 and 2), matching 16-bit training quality on models up to 175B parameters.
~20 min readarXiv:2209.05433✓ verified 2026-10-05
Essential

Distilling the Knowledge in a Neural Network

Geoffrey Hinton, Oriol Vinyals, Jeff Dean · 2015

Defined knowledge distillation: train a small student on a large teacher's softened output probabilities, which carry more information than hard labels.

Problem
Large models and ensembles predict well but are too expensive to deploy widely.
What was new
Raise the softmax temperature so the teacher's probabilities for wrong answers become visible, and train the student to match them at the same temperature.

How to read it: Section 2 explains temperature and soft targets in a page.

~25 min readarXiv:1503.02531✓ verified 2026-10-05
Optional

DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter

Victor Sanh, Lysandre Debut et al. · 2019

A widely used demonstration that distillation during pretraining produces a much cheaper general-purpose language model.

Problem
BERT-sized models were too heavy for many deployments.
What was new
Distill during pretraining with a combined loss: 40% smaller, 60% faster, retaining 97% of BERT's language-understanding score.
~15 min readarXiv:1910.01108✓ verified 2026-10-05
Optional

Learning both Weights and Connections for Efficient Neural Networks

Song Han, Jeff Pool et al. · 2015

The classic train–prune–retrain recipe that showed most connections in a trained network can be removed.

Problem
Networks are larger than they need to be, which costs memory, energy and time on devices.
What was new
Train, remove small-magnitude connections, then retrain the rest. AlexNet went from 61 million to 6.7 million parameters, and VGG-16 from 138 to 10.3 million, without losing accuracy.
~20 min readarXiv:1506.02626✓ verified 2026-10-05
Optional

The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks

Jonathan Frankle, Michael Carbin · 2018

A provocative finding about why pruning works: dense networks contain small subnetworks that could have been trained alone.

Problem
Pruned networks perform well, but training the same sparse architecture from scratch usually doesn't.
What was new
Reset the surviving weights to their original initial values and retrain: winning tickets under 10–20% of the original size matched the full network on MNIST and CIFAR-10.

How to read it: Read the introduction and Section 2. The experiments are on small vision networks; whether it carries over to LLMs is a separate question.

~40 min readarXiv:1803.03635✓ verified 2026-10-05
Optional

SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot

Elias Frantar, Dan Alistarh · 2023

Showed that GPT-scale models can lose half their weights in one pass, without retraining.

Problem
Pruning methods that need retraining are impractical for 100B+ parameter models.
What was new
A one-shot pruning solver that reached 50–60% unstructured sparsity on OPT-175B and BLOOM-176B in under 4.5 hours, with a small perplexity increase.
~40 min readarXiv:2301.00774✓ verified 2026-10-05
Optional

Parameter-Efficient Transfer Learning for NLP

Neil Houlsby, Andrei Giurgiu et al. · 2019

Introduced adapters, small trainable layers inserted into a frozen network: the start of parameter-efficient fine-tuning.

Problem
Fine-tuning a separate copy of a large model for every task is wasteful when tasks arrive one after another.
What was new
Insert small bottleneck layers into each Transformer block and train only those. On GLUE, within 0.4% of full fine-tuning while adding 3.6% parameters per task.

How to read it: Figure 2 shows where the adapters go; that picture is most of the paper.

~25 min readarXiv:1902.00751✓ verified 2026-10-05
Essential

LoRA: Low-Rank Adaptation of Large Language Models

Edward J. Hu, Yelong Shen et al. · 2021

The standard way to adapt a large model cheaply: freeze its weights and learn a small low-rank update for some matrices.

Problem
Full fine-tuning updates and stores every parameter for every task: for GPT-3 175B, a 350 GB checkpoint per task.
What was new
Represent each weight update as the product of two thin matrices (rank r), train only those, and merge them into the weights afterwards, so there is no extra inference latency. For GPT-3 it cut trainable parameters about 10,000-fold and GPU memory about threefold.

How to read it: Section 4 (the method, one page) and Section 7 (why a low rank is enough) are the essentials.

~40 min readarXiv:2106.09685✓ verified 2026-10-05
Essential

QLoRA: Efficient Finetuning of Quantized LLMs

Tim Dettmers, Artidoro Pagnoni et al. · 2023

Combined 4-bit quantization with LoRA so a 65B model can be fine-tuned on one 48 GB GPU, putting large-model fine-tuning within reach of small labs.

Problem
Even with LoRA, a large model's 16-bit frozen weights need several GPUs just to hold them during fine-tuning.
What was new
Freeze the base model in a 4-bit NormalFloat format, backpropagate through it into LoRA adapters, quantize the quantization constants too (double quantization), and page optimizer states to CPU memory during spikes.

How to read it: Section 3 explains NF4, double quantization and paged optimizers in two pages; the block-size arithmetic (0.5 → 0.127 bits per parameter) is worth checking by hand.

~45 min readarXiv:2305.14314✓ verified 2026-10-05
Essential

Fast Inference from Transformers via Speculative Decoding

Yaniv Leviathan, Matan Kalman, Yossi Matias · 2022

Speeds up generation without changing its output distribution: a small model guesses ahead and the large model checks the guesses in parallel.

Problem
Generating K tokens takes K sequential runs of the large model, and each run leaves most of the hardware idle.
What was new
Draft γ tokens with a cheap model, score them with one parallel pass of the target model, and accept or resample with a rule that provably preserves the target's distribution. 2–3× faster on T5-XXL with identical outputs.

How to read it: Section 2.3 (the accept/resample rule) and Equation 1 (expected tokens per pass) are what the lab in this chapter implements.

~35 min readarXiv:2211.17192✓ verified 2026-10-05
Optional

Accelerating Large Language Model Decoding with Speculative Sampling

Charlie Chen, Sebastian Borgeaud et al. · 2023

Independently arrived at the same idea and showed it on a 70B model in a distributed setting.

Problem
Sampling from a large Transformer is memory-bandwidth bound and slow.
What was new
Score short drafts from a faster model in parallel with a modified rejection sampling scheme that preserves the target distribution: a 2–2.5× decoding speed-up on Chinchilla 70B.
~20 min readarXiv:2302.01318✓ verified 2026-10-05
Important

Efficient Memory Management for Large Language Model Serving with PagedAttention

Woosuk Kwon, Zhuohan Li et al. · 2023

Brought operating-system paging to the KV cache (the vLLM server), so many more requests fit in GPU memory at once.

Problem
Serving systems reserved one contiguous KV-cache region per request for its maximum length; only 20–38% of the reserved memory held actual token states.
What was new
Store each request's cache in fixed-size blocks located through a block table, allocated on demand and shareable between requests. Throughput rose 2–4× at the same latency.

How to read it: Section 3 (why existing systems waste memory) and Figure 6 (the block table) are the core.

~45 min readarXiv:2309.06180✓ verified 2026-10-05
Important

Llama 2: Open Foundation and Fine-Tuned Chat Models

Hugo Touvron, Louis Martin et al. · 2023

A widely used open model family; its larger models adopted grouped-query attention for faster inference.

Problem
Open models lagged behind closed chat models, and serving the larger ones was expensive.
What was new
Pretrained models from 7B to 70B with doubled context length and grouped-query attention in the larger models, plus chat versions trained with SFT and RLHF.

How to read it: For this chapter, read Section 2.2 and Appendix A.2.1, which compares multi-head, multi-query and grouped-query attention.

~1 h 30 min readarXiv:2307.09288✓ verified 2026-10-05
Essential

The Llama 3 Herd of Models

Aaron Grattafiori, Abhimanyu Dubey et al. · 2024

The most complete public account of building a frontier-scale model end to end: data pipeline, scaling-law experiments, 16,384-GPU training, failures and all.

Problem
Frontier labs had largely stopped publishing how their models were built.
What was new
A 405B model trained on 15.6T tokens with 3.8 × 10²⁵ FLOPs, with detailed data cleaning, a data mix chosen by scaling experiments, 4D parallelism at 38–43% utilisation, and 466 job interruptions in a 54-day period.

How to read it: It is 90+ pages. For this chapter read Section 3 (pre-training) only: data, scaling laws, infrastructure and the training recipe.

~2 h readarXiv:2407.21783✓ verified 2026-10-04

Watch

What came next?

Chapter 12

Embeddings, RAG & the LLM Application Stack

A model's knowledge is frozen at training time and can't cite its sources.

This chapter is being written.