Concept · Chapter 15: Multimodal AI
Generative Models: GANs, VAEs and Image Tokens
A generative model learns the distribution of its training data well enough to draw new samples from it; for images the main families are GANs, variational autoencoders, autoregressive models over image tokens, and diffusion.
The problem
Recognizing a cat needs one label per image; drawing a new cat means choosing millions of pixel values that are jointly plausible, which no single prediction captures.
The solution
Learn to sample: train a generator against a critic (GAN), compress and decode through a latent space (VAE), predict image tokens one at a time (autoregressive), or learn to reverse a noising process (diffusion).
The consequence
Each family trades sample quality, diversity, training stability and sampling speed differently; diffusion won images by about 2021, and token-based models stay attractive because they share machinery with LLMs.
You should understand first
- Probability and Distributions
- Text as Data
- Vectors
- One-Hot Encoding
- Tokenization
- Tensors and Shapes
- Images as Tensors
- Dot Product
- Convolution
- The Turing Test
- Symbolic AI
- Logic and Rules
- Expert Systems
- Knowledge Representation
- The Knowledge-Acquisition Bottleneck
- From Rules to Learning
- Supervised, Unsupervised and Self-Supervised Learning
- Features, Labels and Tasks
- Loss Functions
- Derivatives and Gradients
- Gradient Descent
- Linear Regression
- Entropy
- Softmax
- Cross-Entropy Loss
- Logistic Regression
- The Perceptron
- Activation Functions
- The Artificial Neuron
- Matrix Multiplication
- Multilayer Perceptron (MLP)
- The Chain Rule
- The Forward Pass
- Computational Graphs and Autodiff
- Backpropagation
- Convolutional Neural Networks
- Embeddings
- Attention
- Self-Attention
- Vision Transformer (ViT)
- Audio and Spectrograms
- Turning Signals into Tokens
- Generative Models: GANs, VAEs and Image Tokens
Why generation is harder than recognition
A classifier answers one question per image. A generator must produce every pixel, and the pixels must agree with one another: a left eye implies a right eye at the matching height. The target is a whole probability distribution over images, and the model has to sample from it, not just score it. Language models solve the same problem for text with autoregressive generation: one token at a time, each conditioned on the ones before.
Four families
GANs. Goodfellow and colleagues trained a generator to maximize the chance that a discriminator mistakes its samples for training data, a two-player game whose ideal solution recovers the data distribution. Established GANs produced sharp faces and scenes through the late 2010s. In practice the game is hard to balance, and generators can collapse onto a few kinds of output (mode collapse).
Variational autoencoders. Kingma and Welling made it possible to train an encoder–decoder with a continuous latent code by stochastic gradients, using the reparameterization trick. Established VAEs train stably and give smooth latent spaces; their samples were historically blurrier. Their encoder–decoder half is now the compressor inside latent diffusion.
Autoregressive over tokens. Turn an image into codebook tokens and train a Transformer to predict them one by one, after the text. DALL·E did exactly this with 256 text tokens followed by 1,024 image tokens and a 12-billion-parameter Transformer. Established The appeal is reuse: the same architecture, training code and serving stack as an LLM, and one model that can interleave text and images.
Diffusion. Learn to remove noise, then generate by denoising pure noise step by step (diffusion models). In 2021 Dhariwal and Nichol reported diffusion models beating the best GANs on ImageNet sample quality while covering the distribution better. Established
Tiny comparison
| Family | Sampling | Main weakness |
|---|---|---|
| GAN | one network call | unstable training, mode collapse |
| VAE | one network call | blurrier samples |
| Autoregressive tokens | one call per token (1,024 for one DALL·E image) | slow, lossy codes |
| Diffusion | one call per step (tens to hundreds) | slow sampling, memorization risk |
Mini experiment
Write the sampling cost of each family for one 256 × 256 image in network calls, using the table. Then ask: if a model generated text and images in one stream, which family would let it share the most machinery with an LLM, and what would it give up?
What to remember
- Discriminative: p(label | image). Generative: samples from p(image), or p(image | text).
- GAN: generator vs discriminator in a minimax game; sharp samples, unstable training, can drop parts of the distribution.
- VAE: encoder to a latent code, decoder back; stable and smooth, historically blurrier samples; its autoencoder lives on inside latent diffusion.
- Autoregressive over image tokens: an LLM-style Transformer predicts codebook indices (DALL·E, Chameleon).
- Diffusion: learn to denoise, sample by iterated denoising; overtook GANs on image benchmarks in 2021.
Key papers
Generative Adversarial Networks
Ian J. Goodfellow, Jean Pouget-Abadie et al. · 2014
The first family of deep generative models to produce convincing images. GANs dominated image synthesis until diffusion models overtook them around 2021.
How to read it: Read the minimax objective and the theoretical result (the optimum recovers the data distribution), then notice how little the paper can promise about reaching that optimum in practice.
Auto-Encoding Variational Bayes
Diederik P Kingma, Max Welling · 2013
The variational autoencoder: an encoder that compresses data to a latent code and a decoder that reconstructs it, trained with a tractable bound. Its descendants compress images for latent diffusion.
How to read it: The key move is the reparameterization trick in Section 2.4: write a random sample as a deterministic function of the parameters plus independent noise, so gradients flow through it.
Neural Discrete Representation Learning
Aaron van den Oord, Oriol Vinyals, Koray Kavukcuoglu · 2017
VQ-VAE turned images, video and speech into sequences of discrete codes from a learned codebook. That is the trick that lets a language-model-style Transformer generate pictures and sound as tokens.
How to read it: Focus on the nearest-neighbour lookup and the straight-through gradient: the encoder never sees the rounding, it receives the decoder's gradient as if no rounding had happened.
Zero-Shot Text-to-Image Generation
Aditya Ramesh, Mikhail Pavlov et al. · 2021
DALL·E treated text-to-image generation as language modelling over one stream of text tokens followed by image tokens, and showed that scale made it work.
How to read it: Section 2 is the recipe in two stages. Count the tokens: 256 + 1,024 per example is why image tokens are expensive context.
Diffusion Models Beat GANs on Image Synthesis
Prafulla Dhariwal, Alex Nichol · 2021
The paper whose title marked the handover from GANs to diffusion, and the origin of guidance as a quality-for-diversity dial.
How to read it: Section 4 is the part to read: one hyperparameter, the scale of the classifier gradients, trades diversity (measured as recall) for fidelity (precision).