Skip to content
Road to Intelligence

Concept · Chapter 11: Inside Modern LLMs

Speculative Decoding

Should knowUnderstand12 minDifficulty

Speculative decoding lets a small draft model guess several tokens ahead and has the large model check all the guesses in one parallel pass, keeping the ones it agrees with by a rule that leaves the output distribution exactly the large model's own.

The problem

Each generated token needs a full, sequential pass of the large model, and each pass leaves most of the GPU's arithmetic idle because decoding is memory-bound.

The solution

Draft γ tokens cheaply, score all of them (plus one more position) with a single target pass, accept draft token x with probability min(1, p(x)/q(x)), and at the first rejection sample a replacement from what's left of p.

The consequence

Generation gets faster, typically 2–3× in the published experiments, with no change to what the model would have sampled. The speed-up depends on how often the draft agrees with the target and how cheap the draft is.

Spare arithmetic

A decode step for one sequence leaves most of a GPU's arithmetic unused (prefill and decode). Scoring five positions at once costs about the same time as scoring one, because the weights are read once either way. Speculative decoding spends that spare capacity on guesses.

Leviathan and colleagues observed that hard language-modelling tasks often include easier steps that a much smaller model can approximate well, and used speculative execution with a new sampling rule to make exact decoding from the large model faster Established.

Draft, then verify

  1. A small draft model qq proposes γ\gamma tokens, one after another (cheap).
  2. The large target model pp scores all γ\gamma positions, plus one more, in a single parallel pass.
  3. Walk through the drafts. Keep draft token xx with probability min⁡(1,p(x)/q(x))\min(1, p(x)/q(x)).
  4. At the first rejection, sample a replacement from the leftover distribution max⁡(0,p−q)\max(0, p - q), normalized, and stop. If every draft is kept, take one more token from the target's extra position.

Each round produces between 1 and γ+1\gamma + 1 tokens for one pass of the target model.

Why the output doesn't change

The accept-or-resample rule is a form of rejection sampling. When the draft over-proposes a token (q>pq > p), it is kept only part of the time; when it under-proposes one (q<pq < p), the shortfall comes back through the resampling step. Both papers prove the procedure samples from exactly the target model's distribution Established, so unlike distillation or quantization, it trades nothing in output quality.

Tiny example. Target p=[0.5,0.25,0.15,0.10]p = [0.5, 0.25, 0.15, 0.10] over four words, draft q=[0.3,0.4,0.2,0.1]q = [0.3, 0.4, 0.2, 0.1]. The draft is accepted with probability ∑min⁡(p,q)=0.3+0.25+0.15+0.1=0.8\sum \min(p, q) = 0.3 + 0.25 + 0.15 + 0.1 = 0.8. Run the rule many times and the words come out in the proportions of pp, not qq: try it in the lab below.

How much faster

If each drafted token is accepted with probability α\alpha (independently, a simplifying assumption the paper makes),

E[tokens per target pass]=1−αγ+11−α,speed-up=1−αγ+1(1−α)(γc+1),\mathbb{E}[\text{tokens per target pass}] = \frac{1 - \alpha^{\gamma+1}}{1 - \alpha}, \qquad \text{speed-up} = \frac{1 - \alpha^{\gamma+1}}{(1 - \alpha)(\gamma c + 1)},

where cc is the cost of a draft step relative to a target step Established. With α=0.8\alpha = 0.8, γ=4\gamma = 4 and c=0.05c = 0.05: 3.36 tokens per pass and a 2.8× speed-up, in theory.

On T5-XXL, speculative decoding ran 2–3× faster than the standard T5X implementation with identical outputs Established, and speculative sampling sped up decoding of the 70B-parameter Chinchilla by 2–2.5× in a distributed setup Established.

Try it · toy model

Draft and Verify

Run speculative decoding's accept-or-resample rule thousands of times and watch the output keep the large model's distribution, then trade acceptance rate, draft length and draft cost for speed.

Understand8 min

What to remember

  • Draft γ tokens with a cheap model; verify all of them in one target pass.
  • Keep draft token x with probability min(1, p(x)/q(x)); on rejection, resample from max(0, p − q), normalized.
  • The output has exactly the target model's distribution.
  • Expected tokens per target pass: (1 − α^(γ+1)) ÷ (1 − α), where α is the acceptance rate.
  • Reported: 2–3× on T5-XXL with identical outputs; 2–2.5× on Chinchilla 70B.

Key papers

Essential

Fast Inference from Transformers via Speculative Decoding

Yaniv Leviathan, Matan Kalman, Yossi Matias · 2022

Speeds up generation without changing its output distribution: a small model guesses ahead and the large model checks the guesses in parallel.

How to read it: Section 2.3 (the accept/resample rule) and Equation 1 (expected tokens per pass) are what the lab in this chapter implements.

~35 min readarXiv:2211.17192✓ verified 2026-10-05

Watch