Skip to content
Road to Intelligence

Concept · Chapter 15: Multimodal AI

Diffusion Models

Must knowKnow well18 minDifficulty

A diffusion model learns to remove noise; to generate, it starts from pure random noise and repeatedly predicts the clean data behind it, stepping a little closer each time, until a sample emerges.

The problem

Generating a whole image in one shot is hard to learn: GANs were unstable, and predicting pixels one by one was slow and awkward for images.

The solution

Destroy data with a known noising process, train a network to undo a little of it at any noise level (usually by predicting the added noise), and generate by running that denoiser from pure noise through many small steps.

The consequence

Training is a stable regression problem and samples are diverse and high quality, but generation costs many network calls, so step counts, samplers and distillation became central engineering questions.

Creating noise is easy

Take a clean data point x0x_0 (an image, or in the lab a point in 2-D) and mix it with Gaussian noise ε\varepsilon:

xt=αˉt x0+1−αˉt ε,ε∼N(0,I).x_t = \sqrt{\bar\alpha_t}\,x_0 + \sqrt{1-\bar\alpha_t}\,\varepsilon, \qquad \varepsilon \sim \mathcal N(0, I).

The schedule αˉt\bar\alpha_t falls from 1 (clean) to nearly 0 (pure noise) as tt goes from 0 to 1. There is nothing to learn here: any noise level can be produced in one line. The idea of learning to reverse a gradual noising process comes from Sohl-Dickstein and colleagues in 2015. Established

Learning to undo it

Train a network εθ(xt,t)\varepsilon_\theta(x_t, t) to recover the noise that was added:

L=Ex0, ε, t ∥ε−εθ(xt,t)∥2.\mathcal L = \mathbb E_{x_0,\,\varepsilon,\,t}\,\big\|\varepsilon - \varepsilon_\theta(x_t, t)\big\|^2 .

That is ordinary supervised regression with labels you make yourself, which is why diffusion trains stably. DDPM (Ho, Jain and Abbeel, 2020) used this simplified objective with T = 1000 noise levels and reached an FID of 3.17 on CIFAR-10. Established Predicting the noise is equivalent to predicting the clean data: x^0=(xt−1−αˉt ε^)/αˉt\hat x_0 = (x_t - \sqrt{1-\bar\alpha_t}\,\hat\varepsilon)/\sqrt{\bar\alpha_t}.

What the best denoiser does. With squared error, the best possible prediction is an average: x^0=E[x0∣xt]\hat x_0 = \mathbb E[x_0 \mid x_t], the mean of all clean data that could have produced this noisy point. Near the data that average is sharp. From pure noise, every clean point is equally plausible, so the best guess is the average of the whole dataset, which is usually not a valid sample at all. This one fact explains why generation needs many steps.

Generating step by step

Start with pure noise and repeat: predict x^0\hat x_0, then move to a slightly lower noise level, keeping the part of xtx_t that the prediction says was noise:

xt′=αˉt′ x^0+1−αˉt′−σ2  ε^+σz.x_{t'} = \sqrt{\bar\alpha_{t'}}\,\hat x_0 + \sqrt{1-\bar\alpha_{t'}-\sigma^2}\;\hat\varepsilon + \sigma z .

With σ=0\sigma = 0 this is the deterministic DDIM sampler; with the DDPM choice of σ\sigma it adds fresh noise zz at each step. DDIM uses the same trained model as DDPM and produced high-quality samples 10× to 50× faster in wall-clock time in its experiments. Established Each early step commits only to coarse structure; later steps, at lower noise, fill in detail.

Tiny example (lab numbers)

The lab uses a known two-dimensional distribution (points on a ring, a plus and four corner clusters), so the ideal denoiser can be computed exactly, with no network:

  • 1 step: every one of the 240 samples lands on the same point, the average of all the data, at the centre, which lies on no shape.
  • 5 steps: fewer than half the samples are on a shape; the rest sit between shapes, where averages of nearby possibilities fall.
  • 25 steps: 96% of samples are on a shape.

The lab's denoiser is perfect, so its errors come only from taking too few steps. A real network adds its own approximation errors on top.

Relatives and practicalities

  • Score-based view. Song and colleagues showed that diffusion and score-based models are discretizations of one stochastic differential equation, with an equivalent deterministic ODE. Established
  • Flow matching trains a network to predict a velocity that carries noise to data along chosen paths; diffusion paths are one special case.
  • Fewer steps. Better samplers, distillation into few-step students, and consistency-style training all attack the cost of many network calls.
  • Memorization. Carlini and colleagues extracted over a thousand training images from state-of-the-art diffusion models, with duplicated training images most at risk. Established The lab shows the extreme case in miniature: its exact denoiser reproduces its known distribution perfectly and can never produce anything outside it.

Mini experiment

In the lab, set the prompt to "No prompt" and the steps to 1, 2, 5, 10 and 25, pressing "Skip to the end" each time. Record the share on a shape. Then switch to DDPM at 25 steps. Explain why one step gives a single dot at the centre, and why the share on a shape rises with steps even though the denoiser never changes.

Why should I care?

As a researcher

Diffusion and its relatives (score matching, flow matching) are the dominant way to generate images, video and audio, and increasingly robot actions; their maths is the shared language of that literature.

As an engineer

Steps, sampler, guidance scale and seed are the knobs of every image and video API; each step is a full network call, which sets latency and cost.

Modern systems that depend on it

  • text-to-image and text-to-video
  • image editing and inpainting
  • latent diffusion and diffusion transformers
  • diffusion policies for robots

Historical context

Before

GANs generated sharp images but trained unstably and dropped parts of the data distribution; autoregressive pixel models were slow.

After

A plain regression objective, 'predict the noise', trains stably at scale, and sampling quality can be traded against the number of denoising steps.

Used today

Stable Diffusion, Imagen, DALL·E 2 and later, and video generators such as Sora are diffusion models; many audio and robot-action generators are too.

What to remember

  • Forward: x_t = √ᾱ_t · x₀ + √(1 − ᾱ_t) · ε. Noise is added in closed form at any level t.
  • Training: show the network x_t and t, ask for ε (equivalently x₀), minimize squared error.
  • Sampling: start at pure noise, predict x₀, step to a slightly lower noise level, repeat.
  • The ideal denoiser outputs the average clean data consistent with x_t, so from pure noise one step gives the average of all the data.
  • DDPM adds fresh noise each step; DDIM can be deterministic and skip steps (10× to 50× faster in its paper).
  • Each step is a network call: step count trades quality for latency.

Key papers

Important

Deep Unsupervised Learning using Nonequilibrium Thermodynamics

Jascha Sohl-Dickstein, Eric A. Weiss et al. · 2015

The origin of diffusion models: slowly destroy structure with noise, then learn the reverse process that restores it.

How to read it: Read it after DDPM. The idea is all here in 2015; what DDPM added five years later was a simpler training objective and image quality that made people notice.

~45 min readarXiv:1503.03585✓ verified 2026-10-07
Essential

Denoising Diffusion Probabilistic Models

Jonathan Ho, Ajay Jain, Pieter Abbeel · 2020

Made diffusion models practical: a simple noise-prediction objective that produced high-quality images. Almost every modern image and video generator descends from it.

How to read it: Algorithms 1 and 2 (training and sampling) fit in ten lines each and are the paper. Read them first, then the derivation of the simplified loss.

~40 min readarXiv:2006.11239✓ verified 2026-10-07
Important

Denoising Diffusion Implicit Models

Jiaming Song, Chenlin Meng, Stefano Ermon · 2020

Showed that a model trained the DDPM way can be sampled deterministically and in far fewer steps, which is why sampler choice and step count became user-facing settings.

How to read it: Equation 12 is the update the lab uses: predict the clean sample, then step to a lower noise level along the predicted noise direction, optionally adding fresh noise.

~40 min readarXiv:2010.02502✓ verified 2026-10-07
Optional

Score-Based Generative Modeling through Stochastic Differential Equations

Yang Song, Jascha Sohl-Dickstein et al. · 2020

Unified diffusion and score-based models as one continuous-time process, which is the language most later diffusion papers use.

How to read it: Read the abstract's first two sentences and Figure 1, then skip to the probability-flow ODE; the predictor–corrector details can wait.

~1 h readarXiv:2011.13456✓ verified 2026-10-07
Important

Diffusion Models Beat GANs on Image Synthesis

Prafulla Dhariwal, Alex Nichol · 2021

The paper whose title marked the handover from GANs to diffusion, and the origin of guidance as a quality-for-diversity dial.

How to read it: Section 4 is the part to read: one hyperparameter, the scale of the classifier gradients, trades diversity (measured as recall) for fidelity (precision).

~45 min readarXiv:2105.05233✓ verified 2026-10-07
Optional

Flow Matching for Generative Modeling

Yaron Lipman, Ricky T. Q. Chen et al. · 2022

Flow matching trains a model to follow a velocity field from noise to data. It generalizes diffusion paths and is used by several recent image and video generators.

How to read it: Read it as 'diffusion with a different path'. The conditional flow-matching loss is the part that makes training as simple as DDPM's.

~50 min readarXiv:2210.02747✓ verified 2026-10-07
Important

Extracting Training Data from Diffusion Models

Nicholas Carlini, Jamie Hayes et al. · 2023

Showed that image generators can reproduce individual training images, which matters for privacy and copyright and for what 'generating new images' means.

How to read it: Look at how 'memorized' is defined before reading the counts, and note the role of duplicated training images.

~40 min readarXiv:2301.13188✓ verified 2026-10-07

Watch