Skip to content
Road to Intelligence

Concept · Chapter 15: Multimodal AI

Latent Diffusion and Text-to-Image

Must knowUnderstand13 minDifficulty

Latent diffusion compresses images with an autoencoder, runs diffusion on the small latent instead of on pixels, conditions each denoising step on a text embedding, and decodes the result back to an image.

The problem

Diffusion on full-resolution pixels spends most of its computation on imperceptible detail and needs hundreds of GPU days to train and many expensive steps to sample.

The solution

Let a pretrained autoencoder handle the pixel-level detail; train the diffusion model in its latent space (several times smaller per side) and inject the prompt through cross-attention.

The consequence

High-resolution text-to-image generation became cheap enough to train outside the largest labs and to run on consumer GPUs, and the same latent recipe carried over to video.

Compress first, then diffuse

Rombach and colleagues observed that pixel-space diffusion models often consume hundreds of GPU days to train, and moved diffusion into the latent space of a pretrained autoencoder. Established The autoencoder handles what they call perceptual compression (textures and pixel detail that a decoder can restore); the diffusion model spends its capacity on semantic content: what is in the picture and where.

The paper compared spatial downsampling factors f from 1 (pixels) to 32 and found small factors trained slowly while large ones lost fidelity; factors of 4 to 8 worked best. Established With f = 8, a 512 × 512 image becomes a 64 × 64 grid of latent positions, 64 times fewer than its pixels.

Text as a condition

A text encoder (CLIP's, or a large language model's encoder) turns the prompt into a sequence of token embeddings. Cross-attention layers inside the denoiser let every latent position attend to those embeddings at every step, the same attention as in Chapter 7 with queries from the image and keys and values from the text. Latent diffusion introduced these cross-attention layers so that text, bounding boxes or other inputs could condition generation. Established Classifier-free guidance then strengthens the prompt's pull.

Other routes exist. DALL·E 2 first generated a CLIP image embedding from the caption, then decoded an image from that embedding with a diffusion model. Established

From U-Nets to Transformers

Early latent diffusion used a convolutional U-Net as the denoiser. Peebles and Xie's Diffusion Transformer (DiT) runs a Transformer over patches of the latent and found that models with more compute per forward pass consistently reached lower FID. Established That turned image generation into another place where the scaling habits of LLMs apply, and it is the architecture OpenAI's Sora report names for video.

Editing

Generation from noise is one use. To edit an existing image:

  • Partial noising. SDEdit adds noise to an input image or sketch and denoises it with the diffusion model; how much noise you add balances faithfulness to the input against realism. Established
  • Inpainting. Denoise only a masked region while keeping the rest fixed to the (noised) original.
  • Instruction editing. InstructPix2Pix trained a diffusion model on edit examples generated by GPT-3 and Stable Diffusion, so it edits from a written instruction in one forward pass per step, without per-image fine-tuning. Established

Tiny example

Several of the latent diffusion paper's own configurations encode a 256 × 256 RGB image (196,608 numbers) with f = 8 into a latent of shape 32 × 32 × 4 (4,096 numbers): 48 times fewer values for the denoiser to process at every step. Fifty denoising steps with guidance are about 100 denoiser calls on that small latent, plus a single decoder call at the end to return to pixels.

Mini experiment

Using the SDEdit idea in the diffusion lab's terms: if you started the reverse process from a point already on the plus, noised to 30% instead of 100%, would a "ring" prompt be able to move it to the ring? What does that suggest about the noise level as an editing-strength knob?

What to remember

  • Encoder: image → latent several times smaller per side; diffusion happens there; decoder: latent → image.
  • Downsampling factors around 4–8 balanced cost and quality in the latent diffusion paper.
  • Text enters through cross-attention from the denoiser to the text encoder's token embeddings.
  • Diffusion Transformers replace the U-Net with a Transformer over latent patches and scale with compute.
  • Editing: noise an existing image part of the way and denoise (SDEdit), or train on edit pairs (InstructPix2Pix).

Key papers

Optional

SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations

Chenlin Meng, Yutong He et al. · 2021

The simplest editing trick for diffusion models: add some noise to an existing image or sketch, then denoise it. How much noise you add sets the balance between faithfulness and realism.

How to read it: Figure 2 says it all: the starting noise level is the one knob, trading how much of the input survives against how realistic the result is.

~30 min readarXiv:2108.01073✓ verified 2026-10-07
Essential

High-Resolution Image Synthesis with Latent Diffusion Models

Robin Rombach, Andreas Blattmann et al. · 2021

Latent diffusion is the design behind Stable Diffusion: compress the image with an autoencoder, run diffusion in the small latent space, and condition on text through cross-attention.

How to read it: Figure 2 (perceptual versus semantic compression) is the argument for the whole design. Then look at Figure 3 for where the cross-attention conditioning enters.

~45 min readarXiv:2112.10752✓ verified 2026-10-07
Optional

Hierarchical Text-Conditional Image Generation with CLIP Latents

Aditya Ramesh, Prafulla Dhariwal et al. · 2022

DALL·E 2 generated images through CLIP's embedding space, a direct demonstration that the shared space CLIP learned can be run backwards.

How to read it: Look at the variations figures: the embedding fixes the meaning and style, and everything the embedding leaves out changes.

~40 min readarXiv:2204.06125✓ verified 2026-10-07
Optional

Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding

Chitwan Saharia, William Chan et al. · 2022

Imagen showed that the text encoder matters most: a large frozen language model (T5) improved fidelity and prompt alignment more than a larger image model did.

How to read it: Figure 4's scaling comparison is the headline. The DrawBench prompts are worth a skim as a list of what text-to-image models found hard in 2022.

~40 min readarXiv:2205.11487✓ verified 2026-10-07
Optional

InstructPix2Pix: Learning to Follow Image Editing Instructions

Tim Brooks, Aleksander Holynski, Alexei A. Efros · 2022

Editing by instruction ('make it winter') with a diffusion model trained on synthetic edit pairs made by two other models.

How to read it: The data pipeline is the contribution. Compare it with Chapter 10's synthetic instruction data: same idea, other modality.

~30 min readarXiv:2211.09800✓ verified 2026-10-07
Important

Scalable Diffusion Models with Transformers

William Peebles, Saining Xie · 2022

The Diffusion Transformer (DiT) replaced the U-Net with a Transformer over latent patches; OpenAI's Sora report describes Sora as a diffusion transformer.

How to read it: The Gflops-versus-FID plot is the argument: the familiar language-model scaling story, now for a denoiser.

~35 min readarXiv:2212.09748✓ verified 2026-10-07