Concept · Chapter 15: Multimodal AI
Latent Diffusion and Text-to-Image
Latent diffusion compresses images with an autoencoder, runs diffusion on the small latent instead of on pixels, conditions each denoising step on a text embedding, and decodes the result back to an image.
The problem
Diffusion on full-resolution pixels spends most of its computation on imperceptible detail and needs hundreds of GPU days to train and many expensive steps to sample.
The solution
Let a pretrained autoencoder handle the pixel-level detail; train the diffusion model in its latent space (several times smaller per side) and inject the prompt through cross-attention.
The consequence
High-resolution text-to-image generation became cheap enough to train outside the largest labs and to run on consumer GPUs, and the same latent recipe carried over to video.
You should understand first
- Probability and Distributions
- Text as Data
- Vectors
- One-Hot Encoding
- Tokenization
- Tensors and Shapes
- Images as Tensors
- Dot Product
- Convolution
- The Turing Test
- Symbolic AI
- Logic and Rules
- Expert Systems
- Knowledge Representation
- The Knowledge-Acquisition Bottleneck
- From Rules to Learning
- Supervised, Unsupervised and Self-Supervised Learning
- Features, Labels and Tasks
- Loss Functions
- Derivatives and Gradients
- Gradient Descent
- Linear Regression
- Entropy
- Softmax
- Cross-Entropy Loss
- Logistic Regression
- The Perceptron
- Activation Functions
- The Artificial Neuron
- Matrix Multiplication
- Multilayer Perceptron (MLP)
- The Chain Rule
- The Forward Pass
- Computational Graphs and Autodiff
- Backpropagation
- Convolutional Neural Networks
- Embeddings
- Attention
- Self-Attention
- Vision Transformer (ViT)
- Audio and Spectrograms
- Turning Signals into Tokens
- Generative Models: GANs, VAEs and Image Tokens
- Diffusion Models
- Classifier-Free Guidance
- Multi-Head Attention
- Causal Masking
- Positional Encoding
- Residual Connections
- Layer Normalization
- Feed-Forward Sublayer (MLP)
- The Transformer Block
- Encoder, Decoder & Encoder–Decoder
- Text Embeddings
- Contrastive Learning
- CLIP and Image–Text Embeddings
- Latent Diffusion and Text-to-Image
Compress first, then diffuse
Rombach and colleagues observed that pixel-space diffusion models often consume hundreds of GPU days to train, and moved diffusion into the latent space of a pretrained autoencoder. Established The autoencoder handles what they call perceptual compression (textures and pixel detail that a decoder can restore); the diffusion model spends its capacity on semantic content: what is in the picture and where.
The paper compared spatial downsampling factors f from 1 (pixels) to 32 and found small factors trained slowly while large ones lost fidelity; factors of 4 to 8 worked best. Established With f = 8, a 512 × 512 image becomes a 64 × 64 grid of latent positions, 64 times fewer than its pixels.
Text as a condition
A text encoder (CLIP's, or a large language model's encoder) turns the prompt into a sequence of token embeddings. Cross-attention layers inside the denoiser let every latent position attend to those embeddings at every step, the same attention as in Chapter 7 with queries from the image and keys and values from the text. Latent diffusion introduced these cross-attention layers so that text, bounding boxes or other inputs could condition generation. Established Classifier-free guidance then strengthens the prompt's pull.
Other routes exist. DALL·E 2 first generated a CLIP image embedding from the caption, then decoded an image from that embedding with a diffusion model. Established
From U-Nets to Transformers
Early latent diffusion used a convolutional U-Net as the denoiser. Peebles and Xie's Diffusion Transformer (DiT) runs a Transformer over patches of the latent and found that models with more compute per forward pass consistently reached lower FID. Established That turned image generation into another place where the scaling habits of LLMs apply, and it is the architecture OpenAI's Sora report names for video.
Editing
Generation from noise is one use. To edit an existing image:
- Partial noising. SDEdit adds noise to an input image or sketch and denoises it with the diffusion model; how much noise you add balances faithfulness to the input against realism. Established
- Inpainting. Denoise only a masked region while keeping the rest fixed to the (noised) original.
- Instruction editing. InstructPix2Pix trained a diffusion model on edit examples generated by GPT-3 and Stable Diffusion, so it edits from a written instruction in one forward pass per step, without per-image fine-tuning. Established
Tiny example
Several of the latent diffusion paper's own configurations encode a 256 × 256 RGB image (196,608 numbers) with f = 8 into a latent of shape 32 × 32 × 4 (4,096 numbers): 48 times fewer values for the denoiser to process at every step. Fifty denoising steps with guidance are about 100 denoiser calls on that small latent, plus a single decoder call at the end to return to pixels.
Mini experiment
Using the SDEdit idea in the diffusion lab's terms: if you started the reverse process from a point already on the plus, noised to 30% instead of 100%, would a "ring" prompt be able to move it to the ring? What does that suggest about the noise level as an editing-strength knob?
What to remember
- Encoder: image → latent several times smaller per side; diffusion happens there; decoder: latent → image.
- Downsampling factors around 4–8 balanced cost and quality in the latent diffusion paper.
- Text enters through cross-attention from the denoiser to the text encoder's token embeddings.
- Diffusion Transformers replace the U-Net with a Transformer over latent patches and scale with compute.
- Editing: noise an existing image part of the way and denoise (SDEdit), or train on edit pairs (InstructPix2Pix).
Key papers
SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations
Chenlin Meng, Yutong He et al. · 2021
The simplest editing trick for diffusion models: add some noise to an existing image or sketch, then denoise it. How much noise you add sets the balance between faithfulness and realism.
How to read it: Figure 2 says it all: the starting noise level is the one knob, trading how much of the input survives against how realistic the result is.
High-Resolution Image Synthesis with Latent Diffusion Models
Robin Rombach, Andreas Blattmann et al. · 2021
Latent diffusion is the design behind Stable Diffusion: compress the image with an autoencoder, run diffusion in the small latent space, and condition on text through cross-attention.
How to read it: Figure 2 (perceptual versus semantic compression) is the argument for the whole design. Then look at Figure 3 for where the cross-attention conditioning enters.
Hierarchical Text-Conditional Image Generation with CLIP Latents
Aditya Ramesh, Prafulla Dhariwal et al. · 2022
DALL·E 2 generated images through CLIP's embedding space, a direct demonstration that the shared space CLIP learned can be run backwards.
How to read it: Look at the variations figures: the embedding fixes the meaning and style, and everything the embedding leaves out changes.
Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding
Chitwan Saharia, William Chan et al. · 2022
Imagen showed that the text encoder matters most: a large frozen language model (T5) improved fidelity and prompt alignment more than a larger image model did.
How to read it: Figure 4's scaling comparison is the headline. The DrawBench prompts are worth a skim as a list of what text-to-image models found hard in 2022.
InstructPix2Pix: Learning to Follow Image Editing Instructions
Tim Brooks, Aleksander Holynski, Alexei A. Efros · 2022
Editing by instruction ('make it winter') with a diffusion model trained on synthetic edit pairs made by two other models.
How to read it: The data pipeline is the contribution. Compare it with Chapter 10's synthetic instruction data: same idea, other modality.
Scalable Diffusion Models with Transformers
William Peebles, Saining Xie · 2022
The Diffusion Transformer (DiT) replaced the U-Net with a Transformer over latent patches; OpenAI's Sora report describes Sora as a diffusion transformer.
How to read it: The Gflops-versus-FID plot is the argument: the familiar language-model scaling story, now for a denoiser.