Skip to content
Road to Intelligence

Concept · Chapter 15: Multimodal AI

Classifier-Free Guidance

Should knowKnow well11 minDifficulty

Classifier-free guidance makes a diffusion model follow its prompt more strongly by running the denoiser twice per step, with and without the prompt, and extrapolating from the unprompted prediction past the prompted one.

The problem

A prompt-conditioned diffusion model often follows its prompt only loosely, producing generic or off-prompt images.

The solution

Train one model both with the prompt and with it randomly dropped. At sampling time compute both predictions and use ε̂ = ε̂(x) + s·(ε̂(x, c) − ε̂(x)) with a scale s > 1.

The consequence

Prompt adherence and image fidelity rise and diversity falls; push the scale too far and samples become distorted and repetitive. Every guided step costs two network evaluations.

From a classifier to no classifier

Dhariwal and Nichol introduced classifier guidance: add the gradient of an image classifier, trained on noisy images, to the diffusion model's prediction, with a scale that trades diversity for fidelity. Established It worked, but needed a second model trained specially for noisy inputs.

Ho and Salimans showed that guidance needs no classifier: jointly train a conditional and an unconditional model (in practice one network, with the condition dropped at random during training) and combine their predictions to get a similar trade-off between sample quality and diversity. Established

The rule

With prompt cc and guidance scale ss:

ε~=ε^(xt)+s (ε^(xt,c)−ε^(xt)).\tilde\varepsilon = \hat\varepsilon(x_t) + s\,\big(\hat\varepsilon(x_t, c) - \hat\varepsilon(x_t)\big).

The difference ε^(xt,c)−ε^(xt)\hat\varepsilon(x_t, c) - \hat\varepsilon(x_t) is the direction the prompt pushes. Scaling it by s>1s > 1 goes further in that direction than the model itself would. In the paper's notation this is (1+w)ε^(xt,c)−w ε^(xt)(1+w)\hat\varepsilon(x_t,c) - w\,\hat\varepsilon(x_t), so w=s−1w = s - 1; tools usually expose ss.

Because x^0\hat x_0 is a linear function of ε^\hat\varepsilon at a fixed xtx_t, guiding the noise prediction and guiding the clean-data prediction are the same thing. The lab guides x^0\hat x_0, which is easier to draw.

Tiny example (lab numbers)

The lab's prompted model is deliberately imperfect: it reads the prompt only half the time. With the prompt "ring", 25 steps and 240 samples:

Guidance scaleOn the ringRing coverage
0 (prompt ignored)33%89%
1 (plain prompted model)61%97%
395%100%

For the ring, guidance mostly helps. Try "plus" and the cost appears: at scale 8 the samples crowd toward the middle of the plus, only 17% of its blobs get a sample, and fewer than half of the samples land on any shape. The guidance paper itself notes that as guidance strength grows, each conditional places probability mass farther from the other classes, and sample diversity decreases. Established

Why it can distort

Guidance is extrapolation. Where the prompted and unprompted predictions differ a lot, a large scale overshoots. The Imagen paper reports that a large guidance weight improves image–text alignment but damages fidelity, producing highly saturated and unnatural images, and introduces dynamic thresholding to allow higher weights. Established (Imagen uses the same convention as this page: a weight of 1 means no extra guidance.) A guidance scale is a dial between "typical of the data" and "unmistakably the prompt", not a quality knob that is always better higher. Interpretation It plays a role similar to a low temperature in text decoding.

Variations

  • Negative prompts replace the unconditional prediction with one conditioned on what to avoid, so the extrapolation pushes away from it.
  • Cost: two network calls per step (often batched together). Distillation can bake guidance into a single-call student.
  • Text encoders: Imagen found that scaling its frozen text encoder improved fidelity and image–text alignment more than scaling the image diffusion model. Established Guidance amplifies whatever the text encoder understood, and cannot add what it missed.

Mini experiment

In the lab, keep "ring" and 25 steps. Set "Model reads the prompt" to 100% and compare scale 1 with scale 3. Then set it to 30%. When is guidance doing real work, and when is it only reducing diversity? Then switch to "plus" and find the smallest scale at which coverage drops below 60%.

What to remember

  • Guided prediction = unprompted + s × (prompted − unprompted).
  • s = 0 ignores the prompt, s = 1 is the plain prompted model, s > 1 extrapolates. Ho and Salimans write the same rule with w = s − 1.
  • Train with the prompt dropped at random so one network provides both predictions.
  • Higher scale: closer to the prompt, less diverse; too high distorts. In image tools, typical values are well above 1.
  • Two network calls per step, so guidance roughly doubles sampling cost.
  • A negative prompt replaces the unprompted prediction with one conditioned on what you want to avoid.

Key papers

Important

Diffusion Models Beat GANs on Image Synthesis

Prafulla Dhariwal, Alex Nichol · 2021

The paper whose title marked the handover from GANs to diffusion, and the origin of guidance as a quality-for-diversity dial.

How to read it: Section 4 is the part to read: one hyperparameter, the scale of the classifier gradients, trades diversity (measured as recall) for fidelity (precision).

~45 min readarXiv:2105.05233✓ verified 2026-10-07
Optional

Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding

Chitwan Saharia, William Chan et al. · 2022

Imagen showed that the text encoder matters most: a large frozen language model (T5) improved fidelity and prompt alignment more than a larger image model did.

How to read it: Figure 4's scaling comparison is the headline. The DrawBench prompts are worth a skim as a list of what text-to-image models found hard in 2022.

~40 min readarXiv:2205.11487✓ verified 2026-10-07
Essential

Classifier-Free Diffusion Guidance

Jonathan Ho, Tim Salimans · 2022

The guidance method used by most text-to-image and text-to-video systems: the 'guidance scale' slider in image tools is this paper.

How to read it: Algorithms 1 and 2 are short. Note the convention: here w = 0 means no guidance; many tools call w + 1 the guidance scale.

~25 min readarXiv:2207.12598✓ verified 2026-10-07

Watch