Concept · Chapter 15: Multimodal AI
Classifier-Free Guidance
Classifier-free guidance makes a diffusion model follow its prompt more strongly by running the denoiser twice per step, with and without the prompt, and extrapolating from the unprompted prediction past the prompted one.
The problem
A prompt-conditioned diffusion model often follows its prompt only loosely, producing generic or off-prompt images.
The solution
Train one model both with the prompt and with it randomly dropped. At sampling time compute both predictions and use ε̂ = ε̂(x) + s·(ε̂(x, c) − ε̂(x)) with a scale s > 1.
The consequence
Prompt adherence and image fidelity rise and diversity falls; push the scale too far and samples become distorted and repetitive. Every guided step costs two network evaluations.
You should understand first
- Probability and Distributions
- Text as Data
- Vectors
- One-Hot Encoding
- Tokenization
- Tensors and Shapes
- Images as Tensors
- Dot Product
- Convolution
- The Turing Test
- Symbolic AI
- Logic and Rules
- Expert Systems
- Knowledge Representation
- The Knowledge-Acquisition Bottleneck
- From Rules to Learning
- Supervised, Unsupervised and Self-Supervised Learning
- Features, Labels and Tasks
- Loss Functions
- Derivatives and Gradients
- Gradient Descent
- Linear Regression
- Entropy
- Softmax
- Cross-Entropy Loss
- Logistic Regression
- The Perceptron
- Activation Functions
- The Artificial Neuron
- Matrix Multiplication
- Multilayer Perceptron (MLP)
- The Chain Rule
- The Forward Pass
- Computational Graphs and Autodiff
- Backpropagation
- Convolutional Neural Networks
- Embeddings
- Attention
- Self-Attention
- Vision Transformer (ViT)
- Audio and Spectrograms
- Turning Signals into Tokens
- Generative Models: GANs, VAEs and Image Tokens
- Diffusion Models
- Classifier-Free Guidance
From a classifier to no classifier
Dhariwal and Nichol introduced classifier guidance: add the gradient of an image classifier, trained on noisy images, to the diffusion model's prediction, with a scale that trades diversity for fidelity. Established It worked, but needed a second model trained specially for noisy inputs.
Ho and Salimans showed that guidance needs no classifier: jointly train a conditional and an unconditional model (in practice one network, with the condition dropped at random during training) and combine their predictions to get a similar trade-off between sample quality and diversity. EstablishedThe rule
With prompt and guidance scale :
The difference is the direction the prompt pushes. Scaling it by goes further in that direction than the model itself would. In the paper's notation this is , so ; tools usually expose .
Because is a linear function of at a fixed , guiding the noise prediction and guiding the clean-data prediction are the same thing. The lab guides , which is easier to draw.
Tiny example (lab numbers)
The lab's prompted model is deliberately imperfect: it reads the prompt only half the time. With the prompt "ring", 25 steps and 240 samples:
| Guidance scale | On the ring | Ring coverage |
|---|---|---|
| 0 (prompt ignored) | 33% | 89% |
| 1 (plain prompted model) | 61% | 97% |
| 3 | 95% | 100% |
For the ring, guidance mostly helps. Try "plus" and the cost appears: at scale 8 the samples crowd toward the middle of the plus, only 17% of its blobs get a sample, and fewer than half of the samples land on any shape. The guidance paper itself notes that as guidance strength grows, each conditional places probability mass farther from the other classes, and sample diversity decreases. Established
Why it can distort
Guidance is extrapolation. Where the prompted and unprompted predictions differ a lot, a large scale overshoots. The Imagen paper reports that a large guidance weight improves image–text alignment but damages fidelity, producing highly saturated and unnatural images, and introduces dynamic thresholding to allow higher weights. Established (Imagen uses the same convention as this page: a weight of 1 means no extra guidance.) A guidance scale is a dial between "typical of the data" and "unmistakably the prompt", not a quality knob that is always better higher. Interpretation It plays a role similar to a low temperature in text decoding.
Variations
- Negative prompts replace the unconditional prediction with one conditioned on what to avoid, so the extrapolation pushes away from it.
- Cost: two network calls per step (often batched together). Distillation can bake guidance into a single-call student.
- Text encoders: Imagen found that scaling its frozen text encoder improved fidelity and image–text alignment more than scaling the image diffusion model. Established Guidance amplifies whatever the text encoder understood, and cannot add what it missed.
Mini experiment
In the lab, keep "ring" and 25 steps. Set "Model reads the prompt" to 100% and compare scale 1 with scale 3. Then set it to 30%. When is guidance doing real work, and when is it only reducing diversity? Then switch to "plus" and find the smallest scale at which coverage drops below 60%.
What to remember
- Guided prediction = unprompted + s × (prompted − unprompted).
- s = 0 ignores the prompt, s = 1 is the plain prompted model, s > 1 extrapolates. Ho and Salimans write the same rule with w = s − 1.
- Train with the prompt dropped at random so one network provides both predictions.
- Higher scale: closer to the prompt, less diverse; too high distorts. In image tools, typical values are well above 1.
- Two network calls per step, so guidance roughly doubles sampling cost.
- A negative prompt replaces the unprompted prediction with one conditioned on what you want to avoid.
Key papers
Diffusion Models Beat GANs on Image Synthesis
Prafulla Dhariwal, Alex Nichol · 2021
The paper whose title marked the handover from GANs to diffusion, and the origin of guidance as a quality-for-diversity dial.
How to read it: Section 4 is the part to read: one hyperparameter, the scale of the classifier gradients, trades diversity (measured as recall) for fidelity (precision).
Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding
Chitwan Saharia, William Chan et al. · 2022
Imagen showed that the text encoder matters most: a large frozen language model (T5) improved fidelity and prompt alignment more than a larger image model did.
How to read it: Figure 4's scaling comparison is the headline. The DrawBench prompts are worth a skim as a list of what text-to-image models found hard in 2022.
Classifier-Free Diffusion Guidance
Jonathan Ho, Tim Salimans · 2022
The guidance method used by most text-to-image and text-to-video systems: the 'guidance scale' slider in image tools is this paper.
How to read it: Algorithms 1 and 2 are short. Note the convention: here w = 0 means no guidance; many tools call w + 1 the guidance scale.
Watch
3Blue1Brown
But how do AI images and videos actually work? | Guest video by Welch Labs
A guest video by Welch Labs: animated, careful and math-first, the clearest single explanation of how CLIP and diffusion fit together in text-to-image models.