Skip to content
Road to Intelligence

Part IV · Systems

Chapter 15

Multimodal AI

Pictures, sound and action, turned into something a model can read and write.

2 h core path13 concepts3 interactivesCore path · 8 optional concepts foldedDeep · all 13 concepts shown in full

In one sentenceMultimodal systems turn images, audio, video and actions into vectors or tokens a Transformer can process alongside text, and generative models such as diffusion run the process in reverse, from noise to pictures.

The problem

The world is not made of text

A colleague sends you a screenshot of a dashboard: four bars labelled J, F, M, A under the title SIGNUPS, with March clearly short. "Why did signups dip in March?"

Every system in the last fourteen chapters would fail before reaching the question. A language model reads tokens. The screenshot is a grid of numbers: a modest 1,000 × 600 screenshot is 1.8 million of them, with no word boundaries, no vocabulary and no obvious order. The same is true of a voice message (16,000 numbers per second of sound), a video (a stack of screenshots), and a robot's next move (a handful of joint angles).

This chapter is about three jobs:

  1. Reading other modalities: turning pictures, sound and video into something a Transformer can attend to, and connecting that to a language model.
  2. Writing them: generating images, audio and video, which turns out to need a different idea from next-token prediction.
  3. Acting: closing the loop from camera to motor command, and predicting how the world will respond.
Most of this chapter reuses machinery you already know: Transformers (Chapter 7), contrastive embeddings (Chapter 12), pretraining and instruction tuning (Chapters 8–10). What is new is how each modality is turned into a sequence, and one new generative idea, diffusion. Interpretation

Everything becomes a sequence

The Vision Transformer from Chapter 5 already showed the trick for images: cut the picture into square patches, flatten each patch, and treat each one as a token. A 224 × 224 image in 16-pixel patches is 14 × 14 = 196 tokens. Audio gets the same treatment through spectrogram frames, and video through patches that extend across a few frames as well as across pixels.

There are two kinds of token you can make from a patch:

  • Continuous: multiply the patch's numbers by a learned matrix and pass the vector on. Nothing is rounded away. Models that only need to read images do this.
  • Discrete: learn a codebook of reference patches and replace each patch with the index of its nearest entry. The picture becomes a list of integers, exactly like text, which means a model can generate it one token at a time. VQ-VAE introduced learned discrete codes for images, video and speech, and DALL·E turned each 256 × 256 image into 32 × 32 = 1,024 tokens, each one of 8,192 codes, modelled after up to 256 text tokens. Established

Discrete codes are lossy compression, and patches cost context. Try both on the dashboard screenshot from the opening, and on a scene and a receipt.

Try it · toy model

Turn a Picture into Tokens

Cut a chart, a scene or a receipt into patches, then send them as raw numbers or as codebook entries. See the token count, the attention cost and which details survive.

Know well7 min

Three things to notice. First, the token count is set by the patch size: 16-pixel patches give the 64 × 64 chart 16 tokens, 4-pixel patches give it 256, and attention compares every pair, 256 versus 65,536. Second, raw patches lose nothing but every token carries hundreds of numbers. Third, with codes, the title and month letters degrade first: in every setting the lab offers, the error on the lettering is higher than on the rest of the picture. Fine detail is where compression bites, which matters for anyone hoping to read a receipt or a chart axis.

Image tokens are processed like text tokens: they go through prefill, occupy the KV cache, and count against the context window. Established A data-engineering way to see it: patch size is a partition size. Smaller partitions keep more detail per row and multiply the rows every downstream join must touch.

Reading pictures

Pictures and words in one space

Tokens give a Transformer something to read. They do not tell it that a patch of orange fur has anything to do with the word "cat". Classic vision models learned that link one fixed label at a time: an ImageNet model knew exactly 1,000 classes, and a new class meant new labelled photos.

Chapter 12 trained a text encoder by making each question pick its own passage out of a batch. CLIP applies the same idea across modalities. Take a batch of web images with their captions, encode each image with an image encoder and each caption with a text encoder, and fill a grid with their cosine similarities. Training pushes the diagonal (true pairs) up and everything else down, in both directions: each image must find its caption, and each caption its image.

CLIP trained this way on 400 million image–text pairs from the internet with batches of 32,768, and then classified images by comparing them with written label descriptions; it matched the original ResNet-50's ImageNet accuracy without using any of the 1.28 million labelled examples that model was trained on. Established

That last step is the payoff: any text becomes a label. To classify a photo as a dog, a cat or a bird, embed "a photo of a dog", "a photo of a cat" and "a photo of a bird", and pick the closest. The label set is chosen after training.

The lab trains a tiny CLIP in your browser, with the real loss and real gradients. Press Train and watch the similarity grid: at first it is noise, then a diagonal appears.

Try it · toy model

Match Pictures to Words

Train a tiny CLIP in your browser with the real contrastive loss. Watch the similarity grid's diagonal form, then classify pictures zero-shot with labels it never trained on.

Know well8 min

On the default seed, after 400 batches, the toy model picks the right caption out of 16 for about 93% of pictures of combinations it trained on, and about 75% for the four colour–shape combinations it never saw together. Asked to choose among bare colour words that never appeared alone in training, it gets every test picture right. That is zero-shot classification in miniature: words learned in one context reused as labels in another. Across twelve seeds, the never-seen score ranges from 55% to 90%, which is a useful warning: generalizing to combinations a model never saw is real but not guaranteed.

The swap test at the bottom of the lab exposes a second limit. Its text encoder adds word vectors, so "a red circle and a blue square" and "a blue circle and a red square" get the same vector. Real contrastive image–text models use Transformer text encoders, yet a 50,000-case benchmark by Yuksekgonul and colleagues found poor relational understanding, attribute-binding mistakes and a severe lack of word-order sensitivity. Established The training signal explains it: if shuffled captions still pick out the right image in a typical batch, the loss never needed word order. Interpretation

Give a language model eyes

CLIP can match pictures and phrases, but it cannot answer "why did signups dip?" That needs a language model. A vision-language model connects the two. There are three common wirings; step through them.

Three ways to give a language model eyes

Example: LLaVA, LLaVA-1.5

  1. Image336 × 336 pixelsinput
  2. Image encoderCLIP ViT, 14-px patchesfrozen
  3. Projectorone matrix, or a 2-layer MLPnew, trained
  4. Language modelreads image vectors as tokenstrained

What the language model’s token stream holds

▦▦▦… 576Whichmonthdipped?
Visual tokens per image
one per patch: 24 × 24 = 576
Trained first
only the projector, on image–caption pairs
Writes images?
no, text out only

Dashed: frozen pretrained part. Accent: new part trained for the connection. A teaching map of three published designs; many systems mix them, and token counts depend on resolution and settings.

LLaVA connected a CLIP image encoder to a language model with a single trainable projection matrix, first training only that matrix on 595K image–caption pairs, then fine-tuning on 158K visual conversations. Established Those conversations were written by text-only GPT-4 from each image's captions and bounding boxes: the teacher never saw a pixel. Flamingo instead kept a vision encoder and a language model frozen and inserted gated cross-attention layers whose gates start at zero, so the combined model initially behaves exactly like the language model. Established Early-fusion models such as Chameleon put image codes and text in one vocabulary and train one model on both from the start; they can write images as well as read them.

The training recipe mirrors Chapter 10: align first, then instruction-tune. So do the failure modes. A model fine-tuned on answers about captions can answer from what images usually contain rather than what this one shows.

Count the cost, too. LLaVA-1.5 used a CLIP encoder at 336 pixels; with 14-pixel patches that is 24 × 24 = 576 visual tokens per image (our arithmetic). Established Four screenshots in a conversation are about 2,300 tokens before you type a word. That is why resamplers that hand the language model a fixed, small number of visual tokens exist.

Reading the small print

Return to the dashboard. To answer the question, the model must read the title, read four bar heights, see that the short one sits above an "M", and know that the third month is March. Each step can fail. The month letters are three pixels wide; whether they survive depends on the input resolution and patch size, exactly what the tokens lab showed.

Document reading started as a pipeline: an OCR engine extracts text and positions, then a language model reasons over it. Donut was an early OCR-free model that read document images directly, motivated by OCR's cost, its rigidity across languages and layouts, and its errors propagating downstream. Established Today's assistants mostly read pixels end to end. That is simpler, and it uses layout and context together, but when a number is wrong there is no OCR output to inspect. Treat extraction as a noisy parser: ask for structured output, validate totals and formats, and keep the source crop next to each value. Interpretation

Charts add arithmetic on top of reading. MMMU's analysis of 150 GPT-4V errors on college-level image questions attributed 35% to perception, the largest share. Established

OptionalReading Documents, Charts and Screens· folded on the core path. Open it, or switch to Deep to show it here.

Listening and speaking

Chapter 5 turned sound into a spectrogram and Whisper into text. Whisper reads 30-second chunks as 80-channel log-mel spectrograms with a 10 ms stride, and a stride-2 convolution halves the frame count before the Transformer encoder. Established That is 1,500 encoder positions per 30 seconds, about 50 tokens per second of audio, which is enough to read speech.

To write audio token by token, a model needs codes that can be turned back into sound. Neural audio codecs such as SoundStream compress audio with a learned encoder, a stack of residual codebooks and a decoder; at 3 kbps SoundStream beat the Opus codec at 12 kbps in listening tests. Established Each extra codebook encodes what the previous ones missed: the tokens lab's codebook, applied over and over to the leftovers.

These pieces changed voice assistants. OpenAI's GPT-4o announcement (May 2024) describes its earlier voice mode as three chained models (transcription, a text model, speech synthesis) with average latencies of 2.8 to 5.4 seconds that could not hear tone, multiple speakers or background noise; GPT-4o was trained end to end across text, vision and audio and averages 320 ms to respond to audio. Established These are the company's own figures. The lesson generalizes: converting to text early is simple and inspectable, and it throws away whatever text cannot say.

OptionalAudio Language Models and Voice· folded on the core path. Open it, or switch to Deep to show it here.

Writing pictures

Running the arrow backwards

Everything so far reads. Now suppose the colleague asks for a chart illustration, or a product photo for a slide. Generating an image is harder than recognizing one: every pixel must agree with every other, and there are countless valid answers.

Language models generate by predicting one token at a time, and that works for image codes too: DALL·E predicted 1,024 image tokens after the caption. But it took many years of other ideas to get high-quality images. GANs trained a generator to fool a discriminator and produced the sharpest images of the late 2010s, while variational autoencoders learned smooth latent spaces with blurrier samples. Established GANs were notoriously hard to train and prone to producing only a few kinds of output. In 2021, Dhariwal and Nichol reported diffusion models beating the best GANs on ImageNet sample quality while covering the distribution better. Established

OptionalGenerative Models: GANs, VAEs and Image Tokens· folded on the core path. Open it, or switch to Deep to show it here.

Sculpting from noise

Diffusion models go back to Sohl-Dickstein and colleagues in 2015: destroy structure slowly with noise, then learn the reverse process. Established The forward direction needs no learning. Mix a clean sample x0x_0 with Gaussian noise ε\varepsilon at any level tt:

xt=αˉt x0+1−αˉt ε.x_t=\sqrt{\bar\alpha_t}\,x_0+\sqrt{1-\bar\alpha_t}\,\varepsilon .

Training asks a network to look at xtx_t and tt and predict the noise that was added, a plain squared-error regression with labels you make yourself. DDPM trained this objective with 1,000 noise levels and reached state-of-the-art image quality on CIFAR-10 in 2020. Established Predicting the noise is the same as predicting the clean data, since one determines the other.

Here is the key fact. With squared error, the best possible prediction is an average: the mean of all clean data that could have produced this noisy input. Close to the data, that average is sharp. From pure noise, every clean sample is equally plausible, so the best single guess is the average of the entire dataset, which is usually not a valid sample at all. Generation therefore proceeds in many small steps: predict the clean data, move a little toward it, look again with less noise, predict again.

The lab runs this process on a known two-dimensional distribution (a ring, a plus and four corner clusters), so the ideal denoiser can be computed exactly. No network is involved, but the sampling rules are the real ones.

Try it · toy model

Sculpt from Noise

Run real DDIM and DDPM sampling with an exact denoiser. Change the steps and the guidance scale, and watch static become a ring, or collapse when pushed too hard.

Know well9 min

Set the steps to 1 and skip to the end: all 240 samples land on one point, the centre, which is the average of all the data and lies on no shape. With 5 steps fewer than half reach a shape, and with 25 steps 96% do. The denoiser is perfect throughout; only the number of steps changes. The DDIM sampler reuses a DDPM-trained model but can skip steps and run deterministically, and produced high-quality samples 10× to 50× faster in its experiments. Established Every step is a full network call, so step count trades quality for latency.

Deep diveThe update rule the lab usesShould know

At noise level tt the denoiser returns x^0\hat x_0, and the implied noise is ε^=(xt−αˉt x^0)/1−αˉt\hat\varepsilon=(x_t-\sqrt{\bar\alpha_t}\,\hat x_0)/\sqrt{1-\bar\alpha_t}. One step to a lower level t′t' is

xt′=αˉt′ x^0+1−αˉt′−σ2  ε^+σz.x_{t'}=\sqrt{\bar\alpha_{t'}}\,\hat x_0+\sqrt{1-\bar\alpha_{t'}-\sigma^2}\;\hat\varepsilon+\sigma z .

With σ=0\sigma=0 this is deterministic DDIM. With σ2=1−αˉt′1−αˉt(1−αˉtαˉt′)\sigma^2=\frac{1-\bar\alpha_{t'}}{1-\bar\alpha_t}\big(1-\frac{\bar\alpha_t}{\bar\alpha_{t'}}\big) it adds fresh noise zz like DDPM. The lab's denoiser is the exact posterior mean for a mixture of 84 small Gaussian blobs: each blob's responsibility comes from Bayes' rule, and each blob contributes its own shrunk estimate. A trained network approximates this function from samples.

Steering with a prompt

Unconditional diffusion draws something from the data. To draw what the prompt asks for, the denoiser also receives the prompt. That alone often follows the prompt loosely. Classifier-free guidance trains one network both with and without the prompt (dropping it at random), and at sampling time extrapolates from the unprompted prediction past the prompted one. Established With guidance scale ss:

ε~=ε^(xt)+s (ε^(xt,c)−ε^(xt)).\tilde\varepsilon=\hat\varepsilon(x_t)+s\,\big(\hat\varepsilon(x_t,c)-\hat\varepsilon(x_t)\big).

Scale 0 ignores the prompt; scale 1 is the plain prompted model; larger scales push further. (The paper writes the same rule with w=s−1w=s-1.) Each guided step costs two network calls.

Go back to the lab. Its prompted model is deliberately imperfect: it reads the prompt only half the time. With the prompt "ring" at scale 1, 61% of the samples land on the ring; at scale 3, 95% do, and they still cover the whole ring. Now choose "plus" and raise the scale to 8. Every sample that lands on a shape is on the plus, but the samples crowd toward its middle, only 17% of the plus gets a sample, and more than half land on no shape at all. The guidance paper notes that stronger guidance moves each class's probability mass away from the other classes and reduces diversity. Established Imagen reported that large guidance weights improve image–text alignment but produce highly saturated, unnatural images. Established

A guidance scale is a dial between "typical of the data" and "unmistakably the prompt", in the same family as a low temperature in text decoding. Higher is not simply better. Interpretation
OptionalClassifier-Free Guidance· folded on the core path. Open it, or switch to Deep to show it here.

Work in a smaller space

Pixel-space diffusion is expensive: most pixels carry detail a decoder could restore cheaply. Latent diffusion moved diffusion into the latent space of a pretrained autoencoder, found downsampling factors of about 4 to 8 per side worked best, and added cross-attention layers so text and other inputs could condition each denoising step. Established That design, released publicly as Stable Diffusion in August 2022, put text-to-image generation on consumer graphics cards.

The text reaches the image through attention you know from Chapter 7: queries from the latent image positions, keys and values from the text encoder's token embeddings. The Diffusion Transformer later replaced the convolutional U-Net denoiser with a Transformer over latent patches, and found that more compute per forward pass consistently improved image quality. Established

The same machinery edits. Noise an existing image part of the way and denoise it under a new prompt, and the starting noise level sets how much of the original survives. Mask a region and denoise only that, and you have inpainting.

Adding time

A video is a stack of images that must agree with each other through time. Generating frames independently gives flicker; nothing ties frame 40 to frame 1. Video diffusion denoises the whole clip at once. OpenAI's February 2024 Sora report describes compressing video in space and time, cutting the latent into spacetime patches that act as Transformer tokens, and training a diffusion transformer on videos and images of varied duration, resolution and aspect ratio; it gave no model or implementation details. Established An image is just a one-frame video, so one model learns from both.

Token counts are the constraint. Doubling duration doubles the tokens; doubling resolution quadruples them; full attention squares the cost. The same report lists Sora's failures: inaccurate physics in basic interactions such as glass shattering, incorrect changes of object state, and incoherence and spontaneously appearing objects in long samples. Established

OptionalVideo Generation· folded on the core path. Open it, or switch to Deep to show it here.

Acting

From pixels to actions

A robot's next action is a short list of numbers, which means it can be written as tokens. RT-2 discretized each continuous robot-action dimension into 256 bins, wrote an action as 8 integers in a vision-language model's existing vocabulary, and co-fine-tuned the model on robot trajectories plus web vision-language data, calling the result a vision-language-action model. Established In about 6,000 real trials, RT-2 generalized better to new objects and followed instructions absent from the robot data, such as placing an object on a particular number or icon. Established

What transfers from the web is mostly recognition and meaning: knowing which object is "the smallest" or what could serve as an improvised hammer. The motions still come from robot demonstrations, which remain scarce. Interpretation Every action changes the next camera image, so errors compound, just as they did for the agents of Chapter 13, except that a wrong action here moves something physical.

OptionalVision-Language-Action Models· folded on the core path. Open it, or switch to Deep to show it here.

Models that predict the next frame

An agent that can predict the consequences of its actions can practise in imagination. Ha and Schmidhuber compressed game frames with a VAE, predicted the next compressed frame with a recurrent network, and trained a tiny controller entirely inside that learned "dream", then transferred it back to the real game. Established They also found the controller exploiting flaws in its own world model, and that making the dream more uncertain made it harder to cheat: Goodhart's law again.

DreamerV3 learned by imagining futures inside a world model with one configuration across more than 150 tasks, and Genie learned playable 2D environments from unlabelled gameplay video by inferring a small set of latent actions. Established Whether large video generators become reliable world models, models you could plan with, is open. Plausible frames are a weaker test than correct predictions when you change the action. Speculative
OptionalWorld Models· folded on the core path. Open it, or switch to Deep to show it here.

From captions to omni models

GANs date from 2014, but most of this chapter arrived in five crowded years: diffusion, patches and shared spaces first, then the bridges into language models, then sound, video and action in the same token stream.

Multimodal AI, 2020–2024Open the full timeline →

Two threads run through it. Representation: every modality was eventually cut into tokens, so one architecture serves them all. Generation: diffusion, rather than next-token prediction, became the main way to produce images and video, while early-fusion models keep testing whether one token stream can do both. Interpretation Dates here are the papers' and announcements' own; later systems are evolving faster than this chapter can track (checked October 7, 2026).

What did it actually see

Back to the dashboard one last time. A vision-language model answers: "Signups dipped in March, likely due to a seasonal slowdown after a February campaign." The first clause might be read from the pixels. The second is a fluent guess the image cannot support.

Li and colleagues found that large vision-language models frequently describe objects that are not in the image, especially objects common in their instruction data or that often co-occur with what is shown, and proposed POPE, which asks yes/no questions about present and carefully chosen absent objects. Established Generators have their own failure list: realistic but off-prompt, on-prompt but repetitive, or not new at all. Carlini and colleagues extracted over a thousand training images from diffusion models. Established

No single number covers these. The diffusion lab already showed why: guidance 8 on "plus" puts every on-shape sample on the plus, which a prompt-alignment metric would love, while covering only 17% of it.

OptionalMultimodal Hallucination and Evaluation· folded on the core path. Open it, or switch to Deep to show it here.

Before trusting a multimodal result, ask:

  1. What reached the model? Resolution, crop, patch size, frames per second, audio sample rate. Detail lost before encoding cannot be recovered by reasoning.
  2. Which claims are visible, and which are priors? Separate what the image shows from what images like it usually contain.
  3. What was checked? Realism, diversity, prompt alignment and originality are different questions with different metrics.
  4. What did it cost? Visual tokens per image, network calls per sample, seconds of latency.
  5. What would change the answer? Crop, rephrase, add an absent-object question, change the action.

You have now followed the road from symbolic rules to models that read, draw, listen and act. The last chapter turns to the question underneath all of them: how do we measure what these systems can do, know when they are wrong, and look inside them?

Concepts in this chapter

Mark each one as you go. Must-know concepts are the core path.

What do I actually need to remember?

  • A modality reaches a Transformer as a sequence: image patches, audio frames, video spacetime patches, action bins.
  • Tokens can be continuous vectors (for understanding) or codebook indices (so a model can generate them like words).
  • Patch size trades detail per token against token count; attention cost grows with the square of the count.
  • CLIP trains two encoders so matching images and captions point the same way; any text then works as a label.
  • Most vision-language models connect a pretrained image encoder to an LLM with a small trained projector.
  • Diffusion learns to denoise; sampling runs from pure noise through many denoising steps.
  • One denoising step from pure noise gives the average of the data: quality needs steps, and each step costs a network call.
  • Guidance pushes samples toward the prompt at the cost of diversity, and too much distorts them.
  • Latent diffusion compresses first; video adds time; robot actions can be written as tokens.
  • Multimodal models describe things that are not there: check what the model actually saw.

You do not need to memorize everything else. This list is the revision sheet.

Key papers

Essential

Learning Transferable Visual Models From Natural Language Supervision

Alec Radford, Jong Wook Kim et al. · 2021 · ICML 2021

CLIP learned a shared space for images and text from web captions — the backbone of much multimodal AI and text-to-image generation.

Problem
Vision models needed large hand-labelled datasets and only recognized fixed label sets.
What was new
Contrastive training on hundreds of millions of image–caption pairs, enabling zero-shot classification from text descriptions.
~1 h readarXiv:2103.00020✓ verified 2026-09-26
Essential

An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Alexey Dosovitskiy, Lucas Beyer et al. · 2020 · ICLR 2021

Showed that a nearly unmodified Transformer, reading an image as a sequence of patches, can match strong CNNs when pretrained on enough data. Vision and language began to share one architecture.

Problem
Transformers dominated language, but vision still relied on convolutions' built-in locality.
What was new
Cut the image into 16×16 patches, embed each patch like a token, add position embeddings and run a standard Transformer encoder.

How to read it: The key result is the comparison across pretraining dataset sizes: with less data CNNs win, with more the ViT catches up.

~40 min readarXiv:2010.11929✓ verified 2026-09-26
Important

Neural Discrete Representation Learning

Aaron van den Oord, Oriol Vinyals, Koray Kavukcuoglu · 2017

VQ-VAE turned images, video and speech into sequences of discrete codes from a learned codebook. That is the trick that lets a language-model-style Transformer generate pictures and sound as tokens.

Problem
Continuous latent codes paired with a powerful autoregressive decoder tended to be ignored (posterior collapse), and continuous codes do not fit token-based models.
What was new
An encoder whose outputs are snapped to the nearest entry of a learned codebook (vector quantization), with a learned autoregressive prior over the resulting codes.

How to read it: Focus on the nearest-neighbour lookup and the straight-through gradient: the encoder never sees the rounding, it receives the decoder's gradient as if no rounding had happened.

~30 min readarXiv:1711.00937✓ verified 2026-10-07
Important

Zero-Shot Text-to-Image Generation

Aditya Ramesh, Mikhail Pavlov et al. · 2021

DALL·E treated text-to-image generation as language modelling over one stream of text tokens followed by image tokens, and showed that scale made it work.

Problem
Text-to-image systems relied on dataset-specific architectures, auxiliary losses and extra labels.
What was new
A discrete VAE compresses each 256×256 image into a 32×32 grid of tokens from 8,192 possible values; a 12-billion-parameter Transformer models up to 256 text tokens followed by the 1,024 image tokens autoregressively.

How to read it: Section 2 is the recipe in two stages. Count the tokens: 256 + 1,024 per example is why image tokens are expensive context.

~40 min readarXiv:2102.12092✓ verified 2026-10-07
Essential

Flamingo: a Visual Language Model for Few-Shot Learning

Jean-Baptiste Alayrac, Jeff Donahue et al. · 2022

Flamingo bridged a frozen vision encoder and a frozen language model, and handled interleaved images and text, so a visual language model could learn new tasks from a few examples in the prompt.

Problem
Vision–language systems needed task-specific fine-tuning with many labelled examples.
What was new
A Perceiver Resampler turns image or video features into a fixed number of visual tokens, which condition the frozen language model through new gated cross-attention layers; trained on web data with interleaved text and images.

How to read it: Figure 4 shows the gated cross-attention block. Its tanh gate starts at zero, so at initialization the model behaves exactly like the frozen language model.

~1 h readarXiv:2204.14198✓ verified 2026-10-07
Essential

Visual Instruction Tuning

Haotian Liu, Chunyuan Li et al. · 2023

LLaVA showed a simple, open recipe for a visual assistant: a CLIP image encoder, a trained projection into an LLM's word-embedding space, and machine-generated visual instruction data.

Problem
Instruction tuning had transformed text models, but there was little instruction-following data for images.
What was new
Use text-only GPT-4 to write 158K image-grounded conversations from captions and boxes; first train only a projection matrix on 595K image–caption pairs, then fine-tune projection and LLM on the instructions.

How to read it: The two-stage training procedure in Section 4 is the recipe. Section 3 shows that the 'teacher' never saw the images, only their captions and bounding boxes.

~35 min readarXiv:2304.08485✓ verified 2026-10-07
Important

Robust Speech Recognition via Large-Scale Weak Supervision

Alec Radford, Jong Wook Kim et al. · 2022

Whisper: an encoder–decoder Transformer trained on 680,000 hours of audio paired with transcripts gathered from the internet. Speech recognition became one more sequence-to-sequence problem solved by scale.

Problem
Speech recognisers trained on curated datasets were brittle on new accents, noise and domains.
What was new
Train one model on a large, noisy, multilingual weakly supervised dataset; it transcribes, translates and identifies language, and generalises well without fine-tuning.
~40 min readarXiv:2212.04356✓ verified 2026-09-26
Important

Generative Adversarial Networks

Ian J. Goodfellow, Jean Pouget-Abadie et al. · 2014

The first family of deep generative models to produce convincing images. GANs dominated image synthesis until diffusion models overtook them around 2021.

Problem
Generative models with explicit likelihoods were hard to train and sample from for high-dimensional data such as images.
What was new
Train a generator against a discriminator in a two-player game: the generator tries to make the discriminator mistake its samples for real data. No Markov chains are needed to train or sample.

How to read it: Read the minimax objective and the theoretical result (the optimum recovers the data distribution), then notice how little the paper can promise about reaching that optimum in practice.

~30 min readarXiv:1406.2661✓ verified 2026-10-07
Essential

Denoising Diffusion Probabilistic Models

Jonathan Ho, Ajay Jain, Pieter Abbeel · 2020

Made diffusion models practical: a simple noise-prediction objective that produced high-quality images. Almost every modern image and video generator descends from it.

Problem
Diffusion models existed since 2015 but had not produced competitive image quality.
What was new
Train a network to predict the noise that was added to an image (a weighted variational bound connected to denoising score matching), then generate by denoising step by step from pure noise; FID 3.17 on CIFAR-10.

How to read it: Algorithms 1 and 2 (training and sampling) fit in ten lines each and are the paper. Read them first, then the derivation of the simplified loss.

~40 min readarXiv:2006.11239✓ verified 2026-10-07
Essential

Classifier-Free Diffusion Guidance

Jonathan Ho, Tim Salimans · 2022

The guidance method used by most text-to-image and text-to-video systems: the 'guidance scale' slider in image tools is this paper.

Problem
Classifier guidance needed a separate classifier trained on noisy images.
What was new
Train one model both with and without the condition (by dropping it at random), then extrapolate between the two noise predictions, (1 + w)·ε(x, c) − w·ε(x), trading diversity for fidelity like classifier guidance does.

How to read it: Algorithms 1 and 2 are short. Note the convention: here w = 0 means no guidance; many tools call w + 1 the guidance scale.

~25 min readarXiv:2207.12598✓ verified 2026-10-07
Essential

High-Resolution Image Synthesis with Latent Diffusion Models

Robin Rombach, Andreas Blattmann et al. · 2021

Latent diffusion is the design behind Stable Diffusion: compress the image with an autoencoder, run diffusion in the small latent space, and condition on text through cross-attention.

Problem
Pixel-space diffusion models cost hundreds of GPU days to train and many expensive network evaluations to sample.
What was new
Diffuse in the latent space of a pretrained autoencoder (downsampling factor f between 1 and 32, with 4–8 working best), and add cross-attention layers so text, boxes or other inputs can condition generation.

How to read it: Figure 2 (perceptual versus semantic compression) is the argument for the whole design. Then look at Figure 3 for where the cross-attention conditioning enters.

~45 min readarXiv:2112.10752✓ verified 2026-10-07
Important

Scalable Diffusion Models with Transformers

William Peebles, Saining Xie · 2022

The Diffusion Transformer (DiT) replaced the U-Net with a Transformer over latent patches; OpenAI's Sora report describes Sora as a diffusion transformer.

Problem
Diffusion models used convolutional U-Nets, leaving the scaling behaviour of Transformers unused.
What was new
A Transformer over patches of the latent image; models with more Gflops (deeper, wider or with more tokens) consistently reached lower FID, and DiT-XL/2 reached FID 2.27 on class-conditional ImageNet 256×256.

How to read it: The Gflops-versus-FID plot is the argument: the familiar language-model scaling story, now for a denoiser.

~35 min readarXiv:2212.09748✓ verified 2026-10-07
Essential

RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Anthony Brohan, Noah Brown et al. · 2023

RT-2 named and demonstrated the vision-language-action model: a VLM fine-tuned to write robot actions as tokens, carrying some web knowledge into control.

Problem
Robot policies trained only on robot data generalized poorly to new objects and instructions.
What was new
Discretize each action dimension into 256 bins, write an action as 8 integers in the model's own token vocabulary, and co-fine-tune a large VLM on robot trajectories plus web vision-language tasks; evaluated in about 6,000 trials.

How to read it: Section 3.2 is the action-as-text trick. Then read the generalization results with their categories (unseen objects, backgrounds, environments) rather than one average.

~45 min readarXiv:2307.15818✓ verified 2026-10-07
Important

World Models

David Ha, Jürgen Schmidhuber · 2018

A clear, small demonstration of the world-model idea: learn a compressed model of an environment, then train a tiny controller inside the model's own imagined rollouts.

Problem
Reinforcement learning from raw pixels needs large policies and many real interactions.
What was new
A VAE compresses each frame, a recurrent network predicts the next compressed frame, and a very small controller acts on those features; the controller can be trained entirely inside the learned model and transferred back.

How to read it: The interactive version at worldmodels.github.io is the best way in. Watch for where the controller exploits flaws in its own dream, and how the authors counter it with extra randomness.

~25 min readarXiv:1803.10122✓ verified 2026-10-07
Important

Evaluating Object Hallucination in Large Vision-Language Models

Yifan Li, Yifan Du et al. · 2023

A systematic look at vision-language models describing objects that are not in the image, and POPE, a simple yes/no probe for it.

Problem
Vision-language models mention objects inconsistent with the image, and caption-based metrics measured this unstably.
What was new
Found that objects frequent in the visual instruction data, or that commonly co-occur with objects in the image, are especially likely to be hallucinated; proposed POPE, which asks 'Is there a ⟨object⟩ in the image?' with random, popular and adversarial absent objects.

How to read it: The three negative-sampling settings are the clever part: an adversarial absent object is one that often appears alongside what is really there.

~30 min readarXiv:2305.10355✓ verified 2026-10-07
Important

Extracting Training Data from Diffusion Models

Nicholas Carlini, Jamie Hayes et al. · 2023

Showed that image generators can reproduce individual training images, which matters for privacy and copyright and for what 'generating new images' means.

Problem
It was unclear whether diffusion models memorize and emit specific training images.
What was new
A generate-and-filter attack extracted over a thousand training examples from state-of-the-art models, from photos of individual people to logos; hundreds of trained models showed how data and modelling choices affect memorization, with duplicated images most at risk.

How to read it: Look at how 'memorized' is defined before reading the counts, and note the role of duplicated training images.

~40 min readarXiv:2301.13188✓ verified 2026-10-07

Watch

37 min

3Blue1Brown

But how do AI images and videos actually work? | Guest video by Welch Labs

A guest video by Welch Labs: animated, careful and math-first, the clearest single explanation of how CLIP and diffusion fit together in text-to-image models.

Covers: CLIP's shared embedding space, DDPM as learning to reverse noise, the vector-field view, DDIM, DALL·E 2, conditioning, classifier-free guidance and negative prompts.

Must know

What came next?

Chapter 16

Evaluation, Reliability, Safety & Interpretability

Impressive demos are not evidence. How do we actually know?

This chapter is being written.