Concept · Chapter 11: Inside Modern LLMs
Speculative Decoding
Speculative decoding lets a small draft model guess several tokens ahead and has the large model check all the guesses in one parallel pass, keeping the ones it agrees with by a rule that leaves the output distribution exactly the large model's own.
The problem
Each generated token needs a full, sequential pass of the large model, and each pass leaves most of the GPU's arithmetic idle because decoding is memory-bound.
The solution
Draft γ tokens cheaply, score all of them (plus one more position) with a single target pass, accept draft token x with probability min(1, p(x)/q(x)), and at the first rejection sample a replacement from what's left of p.
The consequence
Generation gets faster, typically 2–3× in the published experiments, with no change to what the model would have sampled. The speed-up depends on how often the draft agrees with the target and how cheap the draft is.
You should understand first
- Text as Data
- Probability and Distributions
- Conditional Probability and Bayes' Theorem
- Probability of Sequences
- Language Modeling
- Vectors
- Dot Product
- Embeddings
- Attention
- Softmax
- Self-Attention
- Causal Masking
- Entropy
- Loss Functions
- Cross-Entropy Loss
- Autoregressive Next-Token Prediction
- Matrix Multiplication
- Multi-Head Attention
- Positional Encoding
- Residual Connections
- Layer Normalization
- Feed-Forward Sublayer (MLP)
- The Transformer Block
- Prefill, Decode and the Memory Wall
- Decoding: Greedy, Temperature, Top-k, Top-p
- Speculative Decoding
Spare arithmetic
A decode step for one sequence leaves most of a GPU's arithmetic unused (prefill and decode). Scoring five positions at once costs about the same time as scoring one, because the weights are read once either way. Speculative decoding spends that spare capacity on guesses.
Leviathan and colleagues observed that hard language-modelling tasks often include easier steps that a much smaller model can approximate well, and used speculative execution with a new sampling rule to make exact decoding from the large model faster Established.
Draft, then verify
- A small draft model proposes tokens, one after another (cheap).
- The large target model scores all positions, plus one more, in a single parallel pass.
- Walk through the drafts. Keep draft token with probability .
- At the first rejection, sample a replacement from the leftover distribution , normalized, and stop. If every draft is kept, take one more token from the target's extra position.
Each round produces between 1 and tokens for one pass of the target model.
Why the output doesn't change
The accept-or-resample rule is a form of rejection sampling. When the draft over-proposes a token (), it is kept only part of the time; when it under-proposes one (), the shortfall comes back through the resampling step. Both papers prove the procedure samples from exactly the target model's distribution Established, so unlike distillation or quantization, it trades nothing in output quality.
Tiny example. Target over four words, draft . The draft is accepted with probability . Run the rule many times and the words come out in the proportions of , not : try it in the lab below.
How much faster
If each drafted token is accepted with probability (independently, a simplifying assumption the paper makes),
where is the cost of a draft step relative to a target step Established. With , and : 3.36 tokens per pass and a 2.8× speed-up, in theory.
On T5-XXL, speculative decoding ran 2–3× faster than the standard T5X implementation with identical outputs Established, and speculative sampling sped up decoding of the 70B-parameter Chinchilla by 2–2.5× in a distributed setup Established.
Try it · toy model
Run speculative decoding's accept-or-resample rule thousands of times and watch the output keep the large model's distribution, then trade acceptance rate, draft length and draft cost for speed.
What to remember
- Draft γ tokens with a cheap model; verify all of them in one target pass.
- Keep draft token x with probability min(1, p(x)/q(x)); on rejection, resample from max(0, p − q), normalized.
- The output has exactly the target model's distribution.
- Expected tokens per target pass: (1 − α^(γ+1)) ÷ (1 − α), where α is the acceptance rate.
- Reported: 2–3× on T5-XXL with identical outputs; 2–2.5× on Chinchilla 70B.
Key papers
Fast Inference from Transformers via Speculative Decoding
Yaniv Leviathan, Matan Kalman, Yossi Matias · 2022
Speeds up generation without changing its output distribution: a small model guesses ahead and the large model checks the guesses in parallel.
How to read it: Section 2.3 (the accept/resample rule) and Equation 1 (expected tokens per pass) are what the lab in this chapter implements.
Accelerating Large Language Model Decoding with Speculative Sampling
Charlie Chen, Sebastian Borgeaud et al. · 2023
Independently arrived at the same idea and showed it on a 70B model in a distributed setting.
Watch
Stanford Online
Stanford CS336 Language Modeling from Scratch | Spring 2025 | Lecture 10: Inference
The closest single lecture to this chapter: the arithmetic of serving a language model and the main ways to make it cheaper.