Skip to content
Road to Intelligence

Concept · Chapter 15: Multimodal AI

Generative Models: GANs, VAEs and Image Tokens

Should knowUnderstand12 minDifficulty

A generative model learns the distribution of its training data well enough to draw new samples from it; for images the main families are GANs, variational autoencoders, autoregressive models over image tokens, and diffusion.

The problem

Recognizing a cat needs one label per image; drawing a new cat means choosing millions of pixel values that are jointly plausible, which no single prediction captures.

The solution

Learn to sample: train a generator against a critic (GAN), compress and decode through a latent space (VAE), predict image tokens one at a time (autoregressive), or learn to reverse a noising process (diffusion).

The consequence

Each family trades sample quality, diversity, training stability and sampling speed differently; diffusion won images by about 2021, and token-based models stay attractive because they share machinery with LLMs.

Why generation is harder than recognition

A classifier answers one question per image. A generator must produce every pixel, and the pixels must agree with one another: a left eye implies a right eye at the matching height. The target is a whole probability distribution over images, and the model has to sample from it, not just score it. Language models solve the same problem for text with autoregressive generation: one token at a time, each conditioned on the ones before.

Four families

GANs. Goodfellow and colleagues trained a generator to maximize the chance that a discriminator mistakes its samples for training data, a two-player game whose ideal solution recovers the data distribution. Established GANs produced sharp faces and scenes through the late 2010s. In practice the game is hard to balance, and generators can collapse onto a few kinds of output (mode collapse).

Variational autoencoders. Kingma and Welling made it possible to train an encoder–decoder with a continuous latent code by stochastic gradients, using the reparameterization trick. Established VAEs train stably and give smooth latent spaces; their samples were historically blurrier. Their encoder–decoder half is now the compressor inside latent diffusion.

Autoregressive over tokens. Turn an image into codebook tokens and train a Transformer to predict them one by one, after the text. DALL·E did exactly this with 256 text tokens followed by 1,024 image tokens and a 12-billion-parameter Transformer. Established The appeal is reuse: the same architecture, training code and serving stack as an LLM, and one model that can interleave text and images.

Diffusion. Learn to remove noise, then generate by denoising pure noise step by step (diffusion models). In 2021 Dhariwal and Nichol reported diffusion models beating the best GANs on ImageNet sample quality while covering the distribution better. Established

Tiny comparison

FamilySamplingMain weakness
GANone network callunstable training, mode collapse
VAEone network callblurrier samples
Autoregressive tokensone call per token (1,024 for one DALL·E image)slow, lossy codes
Diffusionone call per step (tens to hundreds)slow sampling, memorization risk
None of these is obsolete. GAN-style adversarial losses survive inside audio codecs and image autoencoders, VAEs as compressors, autoregressive tokens in early-fusion multimodal models, and diffusion as the main generator. Interpretation

Mini experiment

Write the sampling cost of each family for one 256 × 256 image in network calls, using the table. Then ask: if a model generated text and images in one stream, which family would let it share the most machinery with an LLM, and what would it give up?

What to remember

  • Discriminative: p(label | image). Generative: samples from p(image), or p(image | text).
  • GAN: generator vs discriminator in a minimax game; sharp samples, unstable training, can drop parts of the distribution.
  • VAE: encoder to a latent code, decoder back; stable and smooth, historically blurrier samples; its autoencoder lives on inside latent diffusion.
  • Autoregressive over image tokens: an LLM-style Transformer predicts codebook indices (DALL·E, Chameleon).
  • Diffusion: learn to denoise, sample by iterated denoising; overtook GANs on image benchmarks in 2021.

Key papers

Important

Generative Adversarial Networks

Ian J. Goodfellow, Jean Pouget-Abadie et al. · 2014

The first family of deep generative models to produce convincing images. GANs dominated image synthesis until diffusion models overtook them around 2021.

How to read it: Read the minimax objective and the theoretical result (the optimum recovers the data distribution), then notice how little the paper can promise about reaching that optimum in practice.

~30 min readarXiv:1406.2661✓ verified 2026-10-07
Important

Auto-Encoding Variational Bayes

Diederik P Kingma, Max Welling · 2013

The variational autoencoder: an encoder that compresses data to a latent code and a decoder that reconstructs it, trained with a tractable bound. Its descendants compress images for latent diffusion.

How to read it: The key move is the reparameterization trick in Section 2.4: write a random sample as a deterministic function of the parameters plus independent noise, so gradients flow through it.

~45 min readarXiv:1312.6114✓ verified 2026-10-07
Important

Neural Discrete Representation Learning

Aaron van den Oord, Oriol Vinyals, Koray Kavukcuoglu · 2017

VQ-VAE turned images, video and speech into sequences of discrete codes from a learned codebook. That is the trick that lets a language-model-style Transformer generate pictures and sound as tokens.

How to read it: Focus on the nearest-neighbour lookup and the straight-through gradient: the encoder never sees the rounding, it receives the decoder's gradient as if no rounding had happened.

~30 min readarXiv:1711.00937✓ verified 2026-10-07
Important

Zero-Shot Text-to-Image Generation

Aditya Ramesh, Mikhail Pavlov et al. · 2021

DALL·E treated text-to-image generation as language modelling over one stream of text tokens followed by image tokens, and showed that scale made it work.

How to read it: Section 2 is the recipe in two stages. Count the tokens: 256 + 1,024 per example is why image tokens are expensive context.

~40 min readarXiv:2102.12092✓ verified 2026-10-07
Important

Diffusion Models Beat GANs on Image Synthesis

Prafulla Dhariwal, Alex Nichol · 2021

The paper whose title marked the handover from GANs to diffusion, and the origin of guidance as a quality-for-diversity dial.

How to read it: Section 4 is the part to read: one hyperparameter, the scale of the classifier gradients, trades diversity (measured as recall) for fidelity (precision).

~45 min readarXiv:2105.05233✓ verified 2026-10-07