Concept · Chapter 15: Multimodal AI
Diffusion Models
A diffusion model learns to remove noise; to generate, it starts from pure random noise and repeatedly predicts the clean data behind it, stepping a little closer each time, until a sample emerges.
The problem
Generating a whole image in one shot is hard to learn: GANs were unstable, and predicting pixels one by one was slow and awkward for images.
The solution
Destroy data with a known noising process, train a network to undo a little of it at any noise level (usually by predicting the added noise), and generate by running that denoiser from pure noise through many small steps.
The consequence
Training is a stable regression problem and samples are diverse and high quality, but generation costs many network calls, so step counts, samplers and distillation became central engineering questions.
You should understand first
- Probability and Distributions
- Text as Data
- Vectors
- One-Hot Encoding
- Tokenization
- Tensors and Shapes
- Images as Tensors
- Dot Product
- Convolution
- The Turing Test
- Symbolic AI
- Logic and Rules
- Expert Systems
- Knowledge Representation
- The Knowledge-Acquisition Bottleneck
- From Rules to Learning
- Supervised, Unsupervised and Self-Supervised Learning
- Features, Labels and Tasks
- Loss Functions
- Derivatives and Gradients
- Gradient Descent
- Linear Regression
- Entropy
- Softmax
- Cross-Entropy Loss
- Logistic Regression
- The Perceptron
- Activation Functions
- The Artificial Neuron
- Matrix Multiplication
- Multilayer Perceptron (MLP)
- The Chain Rule
- The Forward Pass
- Computational Graphs and Autodiff
- Backpropagation
- Convolutional Neural Networks
- Embeddings
- Attention
- Self-Attention
- Vision Transformer (ViT)
- Audio and Spectrograms
- Turning Signals into Tokens
- Generative Models: GANs, VAEs and Image Tokens
- Diffusion Models
Creating noise is easy
Take a clean data point (an image, or in the lab a point in 2-D) and mix it with Gaussian noise :
The schedule falls from 1 (clean) to nearly 0 (pure noise) as goes from 0 to 1. There is nothing to learn here: any noise level can be produced in one line. The idea of learning to reverse a gradual noising process comes from Sohl-Dickstein and colleagues in 2015. Established
Learning to undo it
Train a network to recover the noise that was added:
That is ordinary supervised regression with labels you make yourself, which is why diffusion trains stably. DDPM (Ho, Jain and Abbeel, 2020) used this simplified objective with T = 1000 noise levels and reached an FID of 3.17 on CIFAR-10. Established Predicting the noise is equivalent to predicting the clean data: .
What the best denoiser does. With squared error, the best possible prediction is an average: , the mean of all clean data that could have produced this noisy point. Near the data that average is sharp. From pure noise, every clean point is equally plausible, so the best guess is the average of the whole dataset, which is usually not a valid sample at all. This one fact explains why generation needs many steps.
Generating step by step
Start with pure noise and repeat: predict , then move to a slightly lower noise level, keeping the part of that the prediction says was noise:
With this is the deterministic DDIM sampler; with the DDPM choice of it adds fresh noise at each step. DDIM uses the same trained model as DDPM and produced high-quality samples 10× to 50× faster in wall-clock time in its experiments. Established Each early step commits only to coarse structure; later steps, at lower noise, fill in detail.
Tiny example (lab numbers)
The lab uses a known two-dimensional distribution (points on a ring, a plus and four corner clusters), so the ideal denoiser can be computed exactly, with no network:
- 1 step: every one of the 240 samples lands on the same point, the average of all the data, at the centre, which lies on no shape.
- 5 steps: fewer than half the samples are on a shape; the rest sit between shapes, where averages of nearby possibilities fall.
- 25 steps: 96% of samples are on a shape.
The lab's denoiser is perfect, so its errors come only from taking too few steps. A real network adds its own approximation errors on top.
Relatives and practicalities
- Score-based view. Song and colleagues showed that diffusion and score-based models are discretizations of one stochastic differential equation, with an equivalent deterministic ODE. Established
- Flow matching trains a network to predict a velocity that carries noise to data along chosen paths; diffusion paths are one special case.
- Fewer steps. Better samplers, distillation into few-step students, and consistency-style training all attack the cost of many network calls.
- Memorization. Carlini and colleagues extracted over a thousand training images from state-of-the-art diffusion models, with duplicated training images most at risk. Established The lab shows the extreme case in miniature: its exact denoiser reproduces its known distribution perfectly and can never produce anything outside it.
Mini experiment
In the lab, set the prompt to "No prompt" and the steps to 1, 2, 5, 10 and 25, pressing "Skip to the end" each time. Record the share on a shape. Then switch to DDPM at 25 steps. Explain why one step gives a single dot at the centre, and why the share on a shape rises with steps even though the denoiser never changes.
Why should I care?
As a researcher
Diffusion and its relatives (score matching, flow matching) are the dominant way to generate images, video and audio, and increasingly robot actions; their maths is the shared language of that literature.
As an engineer
Steps, sampler, guidance scale and seed are the knobs of every image and video API; each step is a full network call, which sets latency and cost.
Modern systems that depend on it
- text-to-image and text-to-video
- image editing and inpainting
- latent diffusion and diffusion transformers
- diffusion policies for robots
Historical context
Before
GANs generated sharp images but trained unstably and dropped parts of the data distribution; autoregressive pixel models were slow.
After
A plain regression objective, 'predict the noise', trains stably at scale, and sampling quality can be traded against the number of denoising steps.
Used today
Stable Diffusion, Imagen, DALL·E 2 and later, and video generators such as Sora are diffusion models; many audio and robot-action generators are too.
What to remember
- Forward: x_t = √ᾱ_t · x₀ + √(1 − ᾱ_t) · ε. Noise is added in closed form at any level t.
- Training: show the network x_t and t, ask for ε (equivalently x₀), minimize squared error.
- Sampling: start at pure noise, predict x₀, step to a slightly lower noise level, repeat.
- The ideal denoiser outputs the average clean data consistent with x_t, so from pure noise one step gives the average of all the data.
- DDPM adds fresh noise each step; DDIM can be deterministic and skip steps (10× to 50× faster in its paper).
- Each step is a network call: step count trades quality for latency.
Key papers
Deep Unsupervised Learning using Nonequilibrium Thermodynamics
Jascha Sohl-Dickstein, Eric A. Weiss et al. · 2015
The origin of diffusion models: slowly destroy structure with noise, then learn the reverse process that restores it.
How to read it: Read it after DDPM. The idea is all here in 2015; what DDPM added five years later was a simpler training objective and image quality that made people notice.
Denoising Diffusion Probabilistic Models
Jonathan Ho, Ajay Jain, Pieter Abbeel · 2020
Made diffusion models practical: a simple noise-prediction objective that produced high-quality images. Almost every modern image and video generator descends from it.
How to read it: Algorithms 1 and 2 (training and sampling) fit in ten lines each and are the paper. Read them first, then the derivation of the simplified loss.
Denoising Diffusion Implicit Models
Jiaming Song, Chenlin Meng, Stefano Ermon · 2020
Showed that a model trained the DDPM way can be sampled deterministically and in far fewer steps, which is why sampler choice and step count became user-facing settings.
How to read it: Equation 12 is the update the lab uses: predict the clean sample, then step to a lower noise level along the predicted noise direction, optionally adding fresh noise.
Score-Based Generative Modeling through Stochastic Differential Equations
Yang Song, Jascha Sohl-Dickstein et al. · 2020
Unified diffusion and score-based models as one continuous-time process, which is the language most later diffusion papers use.
How to read it: Read the abstract's first two sentences and Figure 1, then skip to the probability-flow ODE; the predictor–corrector details can wait.
Diffusion Models Beat GANs on Image Synthesis
Prafulla Dhariwal, Alex Nichol · 2021
The paper whose title marked the handover from GANs to diffusion, and the origin of guidance as a quality-for-diversity dial.
How to read it: Section 4 is the part to read: one hyperparameter, the scale of the classifier gradients, trades diversity (measured as recall) for fidelity (precision).
Flow Matching for Generative Modeling
Yaron Lipman, Ricky T. Q. Chen et al. · 2022
Flow matching trains a model to follow a velocity field from noise to data. It generalizes diffusion paths and is used by several recent image and video generators.
How to read it: Read it as 'diffusion with a different path'. The conditional flow-matching loss is the part that makes training as simple as DDPM's.
Extracting Training Data from Diffusion Models
Nicholas Carlini, Jamie Hayes et al. · 2023
Showed that image generators can reproduce individual training images, which matters for privacy and copyright and for what 'generating new images' means.
How to read it: Look at how 'memorized' is defined before reading the counts, and note the role of duplicated training images.
Watch
3Blue1Brown
But how do AI images and videos actually work? | Guest video by Welch Labs
A guest video by Welch Labs: animated, careful and math-first, the clearest single explanation of how CLIP and diffusion fit together in text-to-image models.