Part IV · Systems
Chapter 15
Multimodal AI
Pictures, sound and action, turned into something a model can read and write.
In one sentenceMultimodal systems turn images, audio, video and actions into vectors or tokens a Transformer can process alongside text, and generative models such as diffusion run the process in reverse, from noise to pictures.
The problem
The world is not made of text
A colleague sends you a screenshot of a dashboard: four bars labelled J, F, M, A under the title SIGNUPS, with March clearly short. "Why did signups dip in March?"
Every system in the last fourteen chapters would fail before reaching the question. A language model reads tokens. The screenshot is a grid of numbers: a modest 1,000 × 600 screenshot is 1.8 million of them, with no word boundaries, no vocabulary and no obvious order. The same is true of a voice message (16,000 numbers per second of sound), a video (a stack of screenshots), and a robot's next move (a handful of joint angles).
This chapter is about three jobs:
- Reading other modalities: turning pictures, sound and video into something a Transformer can attend to, and connecting that to a language model.
- Writing them: generating images, audio and video, which turns out to need a different idea from next-token prediction.
- Acting: closing the loop from camera to motor command, and predicting how the world will respond.
Everything becomes a sequence
The Vision Transformer from Chapter 5 already showed the trick for images: cut the picture into square patches, flatten each patch, and treat each one as a token. A 224 × 224 image in 16-pixel patches is 14 × 14 = 196 tokens. Audio gets the same treatment through spectrogram frames, and video through patches that extend across a few frames as well as across pixels.
There are two kinds of token you can make from a patch:
- Continuous: multiply the patch's numbers by a learned matrix and pass the vector on. Nothing is rounded away. Models that only need to read images do this.
- Discrete: learn a codebook of reference patches and replace each patch with the index of its nearest entry. The picture becomes a list of integers, exactly like text, which means a model can generate it one token at a time. VQ-VAE introduced learned discrete codes for images, video and speech, and DALL·E turned each 256 × 256 image into 32 × 32 = 1,024 tokens, each one of 8,192 codes, modelled after up to 256 text tokens. Established
Discrete codes are lossy compression, and patches cost context. Try both on the dashboard screenshot from the opening, and on a scene and a receipt.
Try it · toy model
Cut a chart, a scene or a receipt into patches, then send them as raw numbers or as codebook entries. See the token count, the attention cost and which details survive.
Three things to notice. First, the token count is set by the patch size: 16-pixel patches give the 64 × 64 chart 16 tokens, 4-pixel patches give it 256, and attention compares every pair, 256 versus 65,536. Second, raw patches lose nothing but every token carries hundreds of numbers. Third, with codes, the title and month letters degrade first: in every setting the lab offers, the error on the lettering is higher than on the rest of the picture. Fine detail is where compression bites, which matters for anyone hoping to read a receipt or a chart axis.
Image tokens are processed like text tokens: they go through prefill, occupy the KV cache, and count against the context window. Established A data-engineering way to see it: patch size is a partition size. Smaller partitions keep more detail per row and multiply the rows every downstream join must touch.
Reading pictures
Pictures and words in one space
Tokens give a Transformer something to read. They do not tell it that a patch of orange fur has anything to do with the word "cat". Classic vision models learned that link one fixed label at a time: an ImageNet model knew exactly 1,000 classes, and a new class meant new labelled photos.
Chapter 12 trained a text encoder by making each question pick its own passage out of a batch. CLIP applies the same idea across modalities. Take a batch of web images with their captions, encode each image with an image encoder and each caption with a text encoder, and fill a grid with their cosine similarities. Training pushes the diagonal (true pairs) up and everything else down, in both directions: each image must find its caption, and each caption its image.
CLIP trained this way on 400 million image–text pairs from the internet with batches of 32,768, and then classified images by comparing them with written label descriptions; it matched the original ResNet-50's ImageNet accuracy without using any of the 1.28 million labelled examples that model was trained on. EstablishedThat last step is the payoff: any text becomes a label. To classify a photo as a dog, a cat or a bird, embed "a photo of a dog", "a photo of a cat" and "a photo of a bird", and pick the closest. The label set is chosen after training.
The lab trains a tiny CLIP in your browser, with the real loss and real gradients. Press Train and watch the similarity grid: at first it is noise, then a diagonal appears.
Try it · toy model
Train a tiny CLIP in your browser with the real contrastive loss. Watch the similarity grid's diagonal form, then classify pictures zero-shot with labels it never trained on.
On the default seed, after 400 batches, the toy model picks the right caption out of 16 for about 93% of pictures of combinations it trained on, and about 75% for the four colour–shape combinations it never saw together. Asked to choose among bare colour words that never appeared alone in training, it gets every test picture right. That is zero-shot classification in miniature: words learned in one context reused as labels in another. Across twelve seeds, the never-seen score ranges from 55% to 90%, which is a useful warning: generalizing to combinations a model never saw is real but not guaranteed.
The swap test at the bottom of the lab exposes a second limit. Its text encoder adds word vectors, so "a red circle and a blue square" and "a blue circle and a red square" get the same vector. Real contrastive image–text models use Transformer text encoders, yet a 50,000-case benchmark by Yuksekgonul and colleagues found poor relational understanding, attribute-binding mistakes and a severe lack of word-order sensitivity. Established The training signal explains it: if shuffled captions still pick out the right image in a typical batch, the loss never needed word order. Interpretation
Give a language model eyes
CLIP can match pictures and phrases, but it cannot answer "why did signups dip?" That needs a language model. A vision-language model connects the two. There are three common wirings; step through them.
Example: LLaVA, LLaVA-1.5
- Image336 × 336 pixelsinput
- Image encoderCLIP ViT, 14-px patchesfrozen
- Projectorone matrix, or a 2-layer MLPnew, trained
- Language modelreads image vectors as tokenstrained
What the language model’s token stream holds
- Visual tokens per image
- one per patch: 24 × 24 = 576
- Trained first
- only the projector, on image–caption pairs
- Writes images?
- no, text out only
Dashed: frozen pretrained part. Accent: new part trained for the connection. A teaching map of three published designs; many systems mix them, and token counts depend on resolution and settings.
LLaVA connected a CLIP image encoder to a language model with a single trainable projection matrix, first training only that matrix on 595K image–caption pairs, then fine-tuning on 158K visual conversations. Established Those conversations were written by text-only GPT-4 from each image's captions and bounding boxes: the teacher never saw a pixel. Flamingo instead kept a vision encoder and a language model frozen and inserted gated cross-attention layers whose gates start at zero, so the combined model initially behaves exactly like the language model. Established Early-fusion models such as Chameleon put image codes and text in one vocabulary and train one model on both from the start; they can write images as well as read them.
The training recipe mirrors Chapter 10: align first, then instruction-tune. So do the failure modes. A model fine-tuned on answers about captions can answer from what images usually contain rather than what this one shows.
Count the cost, too. LLaVA-1.5 used a CLIP encoder at 336 pixels; with 14-pixel patches that is 24 × 24 = 576 visual tokens per image (our arithmetic). Established Four screenshots in a conversation are about 2,300 tokens before you type a word. That is why resamplers that hand the language model a fixed, small number of visual tokens exist.
Reading the small print
Return to the dashboard. To answer the question, the model must read the title, read four bar heights, see that the short one sits above an "M", and know that the third month is March. Each step can fail. The month letters are three pixels wide; whether they survive depends on the input resolution and patch size, exactly what the tokens lab showed.
Document reading started as a pipeline: an OCR engine extracts text and positions, then a language model reasons over it. Donut was an early OCR-free model that read document images directly, motivated by OCR's cost, its rigidity across languages and layouts, and its errors propagating downstream. Established Today's assistants mostly read pixels end to end. That is simpler, and it uses layout and context together, but when a number is wrong there is no OCR output to inspect. Treat extraction as a noisy parser: ask for structured output, validate totals and formats, and keep the source crop next to each value. Interpretation
Charts add arithmetic on top of reading. MMMU's analysis of 150 GPT-4V errors on college-level image questions attributed 35% to perception, the largest share. Established
Listening and speaking
Chapter 5 turned sound into a spectrogram and Whisper into text. Whisper reads 30-second chunks as 80-channel log-mel spectrograms with a 10 ms stride, and a stride-2 convolution halves the frame count before the Transformer encoder. Established That is 1,500 encoder positions per 30 seconds, about 50 tokens per second of audio, which is enough to read speech.
To write audio token by token, a model needs codes that can be turned back into sound. Neural audio codecs such as SoundStream compress audio with a learned encoder, a stack of residual codebooks and a decoder; at 3 kbps SoundStream beat the Opus codec at 12 kbps in listening tests. Established Each extra codebook encodes what the previous ones missed: the tokens lab's codebook, applied over and over to the leftovers.
These pieces changed voice assistants. OpenAI's GPT-4o announcement (May 2024) describes its earlier voice mode as three chained models (transcription, a text model, speech synthesis) with average latencies of 2.8 to 5.4 seconds that could not hear tone, multiple speakers or background noise; GPT-4o was trained end to end across text, vision and audio and averages 320 ms to respond to audio. Established These are the company's own figures. The lesson generalizes: converting to text early is simple and inspectable, and it throws away whatever text cannot say.
Writing pictures
Running the arrow backwards
Everything so far reads. Now suppose the colleague asks for a chart illustration, or a product photo for a slide. Generating an image is harder than recognizing one: every pixel must agree with every other, and there are countless valid answers.
Language models generate by predicting one token at a time, and that works for image codes too: DALL·E predicted 1,024 image tokens after the caption. But it took many years of other ideas to get high-quality images. GANs trained a generator to fool a discriminator and produced the sharpest images of the late 2010s, while variational autoencoders learned smooth latent spaces with blurrier samples. Established GANs were notoriously hard to train and prone to producing only a few kinds of output. In 2021, Dhariwal and Nichol reported diffusion models beating the best GANs on ImageNet sample quality while covering the distribution better. Established
Sculpting from noise
Diffusion models go back to Sohl-Dickstein and colleagues in 2015: destroy structure slowly with noise, then learn the reverse process. Established The forward direction needs no learning. Mix a clean sample with Gaussian noise at any level :
Training asks a network to look at and and predict the noise that was added, a plain squared-error regression with labels you make yourself. DDPM trained this objective with 1,000 noise levels and reached state-of-the-art image quality on CIFAR-10 in 2020. Established Predicting the noise is the same as predicting the clean data, since one determines the other.
Here is the key fact. With squared error, the best possible prediction is an average: the mean of all clean data that could have produced this noisy input. Close to the data, that average is sharp. From pure noise, every clean sample is equally plausible, so the best single guess is the average of the entire dataset, which is usually not a valid sample at all. Generation therefore proceeds in many small steps: predict the clean data, move a little toward it, look again with less noise, predict again.
The lab runs this process on a known two-dimensional distribution (a ring, a plus and four corner clusters), so the ideal denoiser can be computed exactly. No network is involved, but the sampling rules are the real ones.
Try it · toy model
Run real DDIM and DDPM sampling with an exact denoiser. Change the steps and the guidance scale, and watch static become a ring, or collapse when pushed too hard.
Set the steps to 1 and skip to the end: all 240 samples land on one point, the centre, which is the average of all the data and lies on no shape. With 5 steps fewer than half reach a shape, and with 25 steps 96% do. The denoiser is perfect throughout; only the number of steps changes. The DDIM sampler reuses a DDPM-trained model but can skip steps and run deterministically, and produced high-quality samples 10× to 50× faster in its experiments. Established Every step is a full network call, so step count trades quality for latency.
Deep diveThe update rule the lab usesShould know
At noise level the denoiser returns , and the implied noise is . One step to a lower level is
With this is deterministic DDIM. With it adds fresh noise like DDPM. The lab's denoiser is the exact posterior mean for a mixture of 84 small Gaussian blobs: each blob's responsibility comes from Bayes' rule, and each blob contributes its own shrunk estimate. A trained network approximates this function from samples.
Steering with a prompt
Unconditional diffusion draws something from the data. To draw what the prompt asks for, the denoiser also receives the prompt. That alone often follows the prompt loosely. Classifier-free guidance trains one network both with and without the prompt (dropping it at random), and at sampling time extrapolates from the unprompted prediction past the prompted one. Established With guidance scale :
Scale 0 ignores the prompt; scale 1 is the plain prompted model; larger scales push further. (The paper writes the same rule with .) Each guided step costs two network calls.
Go back to the lab. Its prompted model is deliberately imperfect: it reads the prompt only half the time. With the prompt "ring" at scale 1, 61% of the samples land on the ring; at scale 3, 95% do, and they still cover the whole ring. Now choose "plus" and raise the scale to 8. Every sample that lands on a shape is on the plus, but the samples crowd toward its middle, only 17% of the plus gets a sample, and more than half land on no shape at all. The guidance paper notes that stronger guidance moves each class's probability mass away from the other classes and reduces diversity. Established Imagen reported that large guidance weights improve image–text alignment but produce highly saturated, unnatural images. Established
A guidance scale is a dial between "typical of the data" and "unmistakably the prompt", in the same family as a low temperature in text decoding. Higher is not simply better. InterpretationWork in a smaller space
Pixel-space diffusion is expensive: most pixels carry detail a decoder could restore cheaply. Latent diffusion moved diffusion into the latent space of a pretrained autoencoder, found downsampling factors of about 4 to 8 per side worked best, and added cross-attention layers so text and other inputs could condition each denoising step. Established That design, released publicly as Stable Diffusion in August 2022, put text-to-image generation on consumer graphics cards.
The text reaches the image through attention you know from Chapter 7: queries from the latent image positions, keys and values from the text encoder's token embeddings. The Diffusion Transformer later replaced the convolutional U-Net denoiser with a Transformer over latent patches, and found that more compute per forward pass consistently improved image quality. Established
The same machinery edits. Noise an existing image part of the way and denoise it under a new prompt, and the starting noise level sets how much of the original survives. Mask a region and denoise only that, and you have inpainting.
Adding time
A video is a stack of images that must agree with each other through time. Generating frames independently gives flicker; nothing ties frame 40 to frame 1. Video diffusion denoises the whole clip at once. OpenAI's February 2024 Sora report describes compressing video in space and time, cutting the latent into spacetime patches that act as Transformer tokens, and training a diffusion transformer on videos and images of varied duration, resolution and aspect ratio; it gave no model or implementation details. Established An image is just a one-frame video, so one model learns from both.
Token counts are the constraint. Doubling duration doubles the tokens; doubling resolution quadruples them; full attention squares the cost. The same report lists Sora's failures: inaccurate physics in basic interactions such as glass shattering, incorrect changes of object state, and incoherence and spontaneously appearing objects in long samples. Established
Acting
From pixels to actions
A robot's next action is a short list of numbers, which means it can be written as tokens. RT-2 discretized each continuous robot-action dimension into 256 bins, wrote an action as 8 integers in a vision-language model's existing vocabulary, and co-fine-tuned the model on robot trajectories plus web vision-language data, calling the result a vision-language-action model. Established In about 6,000 real trials, RT-2 generalized better to new objects and followed instructions absent from the robot data, such as placing an object on a particular number or icon. Established
What transfers from the web is mostly recognition and meaning: knowing which object is "the smallest" or what could serve as an improvised hammer. The motions still come from robot demonstrations, which remain scarce. Interpretation Every action changes the next camera image, so errors compound, just as they did for the agents of Chapter 13, except that a wrong action here moves something physical.
Models that predict the next frame
An agent that can predict the consequences of its actions can practise in imagination. Ha and Schmidhuber compressed game frames with a VAE, predicted the next compressed frame with a recurrent network, and trained a tiny controller entirely inside that learned "dream", then transferred it back to the real game. Established They also found the controller exploiting flaws in its own world model, and that making the dream more uncertain made it harder to cheat: Goodhart's law again.
DreamerV3 learned by imagining futures inside a world model with one configuration across more than 150 tasks, and Genie learned playable 2D environments from unlabelled gameplay video by inferring a small set of latent actions. Established Whether large video generators become reliable world models, models you could plan with, is open. Plausible frames are a weaker test than correct predictions when you change the action. SpeculativeFrom captions to omni models
GANs date from 2014, but most of this chapter arrived in five crowded years: diffusion, patches and shared spaces first, then the bridges into language models, then sound, video and action in the same token stream.
Two threads run through it. Representation: every modality was eventually cut into tokens, so one architecture serves them all. Generation: diffusion, rather than next-token prediction, became the main way to produce images and video, while early-fusion models keep testing whether one token stream can do both. Interpretation Dates here are the papers' and announcements' own; later systems are evolving faster than this chapter can track (checked October 7, 2026).
What did it actually see
Back to the dashboard one last time. A vision-language model answers: "Signups dipped in March, likely due to a seasonal slowdown after a February campaign." The first clause might be read from the pixels. The second is a fluent guess the image cannot support.
Li and colleagues found that large vision-language models frequently describe objects that are not in the image, especially objects common in their instruction data or that often co-occur with what is shown, and proposed POPE, which asks yes/no questions about present and carefully chosen absent objects. Established Generators have their own failure list: realistic but off-prompt, on-prompt but repetitive, or not new at all. Carlini and colleagues extracted over a thousand training images from diffusion models. Established
No single number covers these. The diffusion lab already showed why: guidance 8 on "plus" puts every on-shape sample on the plus, which a prompt-alignment metric would love, while covering only 17% of it.
Before trusting a multimodal result, ask:
- What reached the model? Resolution, crop, patch size, frames per second, audio sample rate. Detail lost before encoding cannot be recovered by reasoning.
- Which claims are visible, and which are priors? Separate what the image shows from what images like it usually contain.
- What was checked? Realism, diversity, prompt alignment and originality are different questions with different metrics.
- What did it cost? Visual tokens per image, network calls per sample, seconds of latency.
- What would change the answer? Crop, rephrase, add an absent-object question, change the action.
You have now followed the road from symbolic rules to models that read, draw, listen and act. The last chapter turns to the question underneath all of them: how do we measure what these systems can do, know when they are wrong, and look inside them?
Concepts in this chapter
Mark each one as you go. Must-know concepts are the core path.
- CLIP and Image–Text EmbeddingsCLIP trains an image encoder and a text encoder together so that a picture and its caption land on nearby vectors, which lets any written phrase act as a label and gives later models a shared space for pictures and words.Know wellMust know
- Diffusion ModelsA diffusion model learns to remove noise; to generate, it starts from pure random noise and repeatedly predicts the clean data behind it, stepping a little closer each time, until a sample emerges.Know wellMust know
- Latent Diffusion and Text-to-ImageLatent diffusion compresses images with an autoencoder, runs diffusion on the small latent instead of on pixels, conditions each denoising step on a text embedding, and decodes the result back to an image.UnderstandMust know
- Turning Signals into TokensBefore a Transformer can use an image, a sound or a video, the signal is cut into pieces (patches, frames, spacetime blocks) and each piece becomes one token: either a continuous vector or the index of an entry in a learned codebook.Know wellMust know
- Vision-Language ModelsA vision-language model lets a language model read images by encoding each image into vectors and feeding them into the model alongside text tokens, usually through a small trained connector between a pretrained image encoder and a pretrained LLM.Know wellMust know
- Audio Language Models and VoiceAudio language models treat sound as a token sequence, using spectrogram features to listen and neural-codec codes to speak, so one model can hear speech, music and tone and answer in audio without converting everything to text.UnderstandShould know
- Classifier-Free GuidanceClassifier-free guidance makes a diffusion model follow its prompt more strongly by running the denoiser twice per step, with and without the prompt, and extrapolating from the unprompted prediction past the prompted one.Know wellShould know
- Reading Documents, Charts and ScreensDocument understanding means extracting text, numbers and structure from images of pages, charts, forms and screens, either with a separate OCR step or, increasingly, by a vision-language model reading the pixels directly.UnderstandShould know
- Generative Models: GANs, VAEs and Image TokensA generative model learns the distribution of its training data well enough to draw new samples from it; for images the main families are GANs, variational autoencoders, autoregressive models over image tokens, and diffusion.UnderstandShould know
- Multimodal Hallucination and EvaluationEvaluating multimodal models means checking both directions separately: whether a model that reads images reports only what is there, and whether a model that makes images produces realistic, varied, on-prompt and original outputs.UnderstandShould know
- Video GenerationVideo generators extend image diffusion through time: they compress video into a latent, cut it into spacetime patches, and denoise all frames together so motion stays consistent.UnderstandShould know
- Vision-Language-Action ModelsA vision-language-action model is a vision-language model fine-tuned to output robot actions, often written as discretized tokens, from camera images and a natural-language instruction.UnderstandShould know
- World ModelsA world model is a learned model that predicts how an environment will change in response to actions, so that an agent can plan or practise inside its predictions instead of only in the real world.UnderstandFrontier
What do I actually need to remember?
- A modality reaches a Transformer as a sequence: image patches, audio frames, video spacetime patches, action bins.
- Tokens can be continuous vectors (for understanding) or codebook indices (so a model can generate them like words).
- Patch size trades detail per token against token count; attention cost grows with the square of the count.
- CLIP trains two encoders so matching images and captions point the same way; any text then works as a label.
- Most vision-language models connect a pretrained image encoder to an LLM with a small trained projector.
- Diffusion learns to denoise; sampling runs from pure noise through many denoising steps.
- One denoising step from pure noise gives the average of the data: quality needs steps, and each step costs a network call.
- Guidance pushes samples toward the prompt at the cost of diversity, and too much distorts them.
- Latent diffusion compresses first; video adds time; robot actions can be written as tokens.
- Multimodal models describe things that are not there: check what the model actually saw.
You do not need to memorize everything else. This list is the revision sheet.
Key papers
Learning Transferable Visual Models From Natural Language Supervision
Alec Radford, Jong Wook Kim et al. · 2021 · ICML 2021
CLIP learned a shared space for images and text from web captions — the backbone of much multimodal AI and text-to-image generation.
- Problem
- Vision models needed large hand-labelled datasets and only recognized fixed label sets.
- What was new
- Contrastive training on hundreds of millions of image–caption pairs, enabling zero-shot classification from text descriptions.
- Built on
- Attention Is All You Need
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Alexey Dosovitskiy, Lucas Beyer et al. · 2020 · ICLR 2021
Showed that a nearly unmodified Transformer, reading an image as a sequence of patches, can match strong CNNs when pretrained on enough data. Vision and language began to share one architecture.
- Problem
- Transformers dominated language, but vision still relied on convolutions' built-in locality.
- What was new
- Cut the image into 16×16 patches, embed each patch like a token, add position embeddings and run a standard Transformer encoder.
How to read it: The key result is the comparison across pretraining dataset sizes: with less data CNNs win, with more the ViT catches up.
Neural Discrete Representation Learning
Aaron van den Oord, Oriol Vinyals, Koray Kavukcuoglu · 2017
VQ-VAE turned images, video and speech into sequences of discrete codes from a learned codebook. That is the trick that lets a language-model-style Transformer generate pictures and sound as tokens.
- Problem
- Continuous latent codes paired with a powerful autoregressive decoder tended to be ignored (posterior collapse), and continuous codes do not fit token-based models.
- What was new
- An encoder whose outputs are snapped to the nearest entry of a learned codebook (vector quantization), with a learned autoregressive prior over the resulting codes.
- Built on
- Auto-Encoding Variational Bayes
How to read it: Focus on the nearest-neighbour lookup and the straight-through gradient: the encoder never sees the rounding, it receives the decoder's gradient as if no rounding had happened.
Zero-Shot Text-to-Image Generation
Aditya Ramesh, Mikhail Pavlov et al. · 2021
DALL·E treated text-to-image generation as language modelling over one stream of text tokens followed by image tokens, and showed that scale made it work.
- Problem
- Text-to-image systems relied on dataset-specific architectures, auxiliary losses and extra labels.
- What was new
- A discrete VAE compresses each 256×256 image into a 32×32 grid of tokens from 8,192 possible values; a 12-billion-parameter Transformer models up to 256 text tokens followed by the 1,024 image tokens autoregressively.
How to read it: Section 2 is the recipe in two stages. Count the tokens: 256 + 1,024 per example is why image tokens are expensive context.
Flamingo: a Visual Language Model for Few-Shot Learning
Jean-Baptiste Alayrac, Jeff Donahue et al. · 2022
Flamingo bridged a frozen vision encoder and a frozen language model, and handled interleaved images and text, so a visual language model could learn new tasks from a few examples in the prompt.
- Problem
- Vision–language systems needed task-specific fine-tuning with many labelled examples.
- What was new
- A Perceiver Resampler turns image or video features into a fixed number of visual tokens, which condition the frozen language model through new gated cross-attention layers; trained on web data with interleaved text and images.
How to read it: Figure 4 shows the gated cross-attention block. Its tanh gate starts at zero, so at initialization the model behaves exactly like the frozen language model.
Visual Instruction Tuning
Haotian Liu, Chunyuan Li et al. · 2023
LLaVA showed a simple, open recipe for a visual assistant: a CLIP image encoder, a trained projection into an LLM's word-embedding space, and machine-generated visual instruction data.
- Problem
- Instruction tuning had transformed text models, but there was little instruction-following data for images.
- What was new
- Use text-only GPT-4 to write 158K image-grounded conversations from captions and boxes; first train only a projection matrix on 595K image–caption pairs, then fine-tune projection and LLM on the instructions.
How to read it: The two-stage training procedure in Section 4 is the recipe. Section 3 shows that the 'teacher' never saw the images, only their captions and bounding boxes.
Robust Speech Recognition via Large-Scale Weak Supervision
Alec Radford, Jong Wook Kim et al. · 2022
Whisper: an encoder–decoder Transformer trained on 680,000 hours of audio paired with transcripts gathered from the internet. Speech recognition became one more sequence-to-sequence problem solved by scale.
- Problem
- Speech recognisers trained on curated datasets were brittle on new accents, noise and domains.
- What was new
- Train one model on a large, noisy, multilingual weakly supervised dataset; it transcribes, translates and identifies language, and generalises well without fine-tuning.
- Built on
- Attention Is All You Need
Generative Adversarial Networks
Ian J. Goodfellow, Jean Pouget-Abadie et al. · 2014
The first family of deep generative models to produce convincing images. GANs dominated image synthesis until diffusion models overtook them around 2021.
- Problem
- Generative models with explicit likelihoods were hard to train and sample from for high-dimensional data such as images.
- What was new
- Train a generator against a discriminator in a two-player game: the generator tries to make the discriminator mistake its samples for real data. No Markov chains are needed to train or sample.
How to read it: Read the minimax objective and the theoretical result (the optimum recovers the data distribution), then notice how little the paper can promise about reaching that optimum in practice.
Denoising Diffusion Probabilistic Models
Jonathan Ho, Ajay Jain, Pieter Abbeel · 2020
Made diffusion models practical: a simple noise-prediction objective that produced high-quality images. Almost every modern image and video generator descends from it.
- Problem
- Diffusion models existed since 2015 but had not produced competitive image quality.
- What was new
- Train a network to predict the noise that was added to an image (a weighted variational bound connected to denoising score matching), then generate by denoising step by step from pure noise; FID 3.17 on CIFAR-10.
How to read it: Algorithms 1 and 2 (training and sampling) fit in ten lines each and are the paper. Read them first, then the derivation of the simplified loss.
Classifier-Free Diffusion Guidance
Jonathan Ho, Tim Salimans · 2022
The guidance method used by most text-to-image and text-to-video systems: the 'guidance scale' slider in image tools is this paper.
- Problem
- Classifier guidance needed a separate classifier trained on noisy images.
- What was new
- Train one model both with and without the condition (by dropping it at random), then extrapolate between the two noise predictions, (1 + w)·ε(x, c) − w·ε(x), trading diversity for fidelity like classifier guidance does.
How to read it: Algorithms 1 and 2 are short. Note the convention: here w = 0 means no guidance; many tools call w + 1 the guidance scale.
High-Resolution Image Synthesis with Latent Diffusion Models
Robin Rombach, Andreas Blattmann et al. · 2021
Latent diffusion is the design behind Stable Diffusion: compress the image with an autoencoder, run diffusion in the small latent space, and condition on text through cross-attention.
- Problem
- Pixel-space diffusion models cost hundreds of GPU days to train and many expensive network evaluations to sample.
- What was new
- Diffuse in the latent space of a pretrained autoencoder (downsampling factor f between 1 and 32, with 4–8 working best), and add cross-attention layers so text, boxes or other inputs can condition generation.
How to read it: Figure 2 (perceptual versus semantic compression) is the argument for the whole design. Then look at Figure 3 for where the cross-attention conditioning enters.
Scalable Diffusion Models with Transformers
William Peebles, Saining Xie · 2022
The Diffusion Transformer (DiT) replaced the U-Net with a Transformer over latent patches; OpenAI's Sora report describes Sora as a diffusion transformer.
- Problem
- Diffusion models used convolutional U-Nets, leaving the scaling behaviour of Transformers unused.
- What was new
- A Transformer over patches of the latent image; models with more Gflops (deeper, wider or with more tokens) consistently reached lower FID, and DiT-XL/2 reached FID 2.27 on class-conditional ImageNet 256×256.
How to read it: The Gflops-versus-FID plot is the argument: the familiar language-model scaling story, now for a denoiser.
RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
Anthony Brohan, Noah Brown et al. · 2023
RT-2 named and demonstrated the vision-language-action model: a VLM fine-tuned to write robot actions as tokens, carrying some web knowledge into control.
- Problem
- Robot policies trained only on robot data generalized poorly to new objects and instructions.
- What was new
- Discretize each action dimension into 256 bins, write an action as 8 integers in the model's own token vocabulary, and co-fine-tune a large VLM on robot trajectories plus web vision-language tasks; evaluated in about 6,000 trials.
How to read it: Section 3.2 is the action-as-text trick. Then read the generalization results with their categories (unseen objects, backgrounds, environments) rather than one average.
World Models
David Ha, Jürgen Schmidhuber · 2018
A clear, small demonstration of the world-model idea: learn a compressed model of an environment, then train a tiny controller inside the model's own imagined rollouts.
- Problem
- Reinforcement learning from raw pixels needs large policies and many real interactions.
- What was new
- A VAE compresses each frame, a recurrent network predicts the next compressed frame, and a very small controller acts on those features; the controller can be trained entirely inside the learned model and transferred back.
- Built on
- Auto-Encoding Variational Bayes
How to read it: The interactive version at worldmodels.github.io is the best way in. Watch for where the controller exploits flaws in its own dream, and how the authors counter it with extra randomness.
Evaluating Object Hallucination in Large Vision-Language Models
Yifan Li, Yifan Du et al. · 2023
A systematic look at vision-language models describing objects that are not in the image, and POPE, a simple yes/no probe for it.
- Problem
- Vision-language models mention objects inconsistent with the image, and caption-based metrics measured this unstably.
- What was new
- Found that objects frequent in the visual instruction data, or that commonly co-occur with objects in the image, are especially likely to be hallucinated; proposed POPE, which asks 'Is there a ⟨object⟩ in the image?' with random, popular and adversarial absent objects.
- Built on
- Visual Instruction Tuning
How to read it: The three negative-sampling settings are the clever part: an adversarial absent object is one that often appears alongside what is really there.
Extracting Training Data from Diffusion Models
Nicholas Carlini, Jamie Hayes et al. · 2023
Showed that image generators can reproduce individual training images, which matters for privacy and copyright and for what 'generating new images' means.
- Problem
- It was unclear whether diffusion models memorize and emit specific training images.
- What was new
- A generate-and-filter attack extracted over a thousand training examples from state-of-the-art models, from photos of individual people to logos; hundreds of trained models showed how data and modelling choices affect memorization, with duplicated images most at risk.
How to read it: Look at how 'memorized' is defined before reading the counts, and note the role of duplicated training images.
Watch
3Blue1Brown
But how do AI images and videos actually work? | Guest video by Welch Labs
A guest video by Welch Labs: animated, careful and math-first, the clearest single explanation of how CLIP and diffusion fit together in text-to-image models.
Covers: CLIP's shared embedding space, DDPM as learning to reverse noise, the vector-field view, DDIM, DALL·E 2, conditioning, classifier-free guidance and negative prompts.
What came next?
Chapter 16
Evaluation, Reliability, Safety & Interpretability
Impressive demos are not evidence. How do we actually know?
This chapter is being written.