Concept · Chapter 15: Multimodal AI
Turning Signals into Tokens
Before a Transformer can use an image, a sound or a video, the signal is cut into pieces (patches, frames, spacetime blocks) and each piece becomes one token: either a continuous vector or the index of an entry in a learned codebook.
The problem
A Transformer reads a sequence of vectors, but a photo is a grid of millions of numbers, audio is tens of thousands of samples per second, and video is both at once.
The solution
Cut the signal into regular pieces, embed each piece as one vector (continuous tokens), or snap it to the nearest entry of a learned codebook and use the entry's number (discrete tokens).
The consequence
Every modality becomes a sequence, so the machinery of Chapters 7–11 applies. The price is token count: images and video are expensive context, and discrete codes lose detail.
You should understand first
- Text as Data
- Vectors
- One-Hot Encoding
- Tokenization
- Tensors and Shapes
- Images as Tensors
- Dot Product
- Convolution
- The Turing Test
- Symbolic AI
- Logic and Rules
- Expert Systems
- Knowledge Representation
- The Knowledge-Acquisition Bottleneck
- From Rules to Learning
- Supervised, Unsupervised and Self-Supervised Learning
- Features, Labels and Tasks
- Loss Functions
- Derivatives and Gradients
- Gradient Descent
- Linear Regression
- Probability and Distributions
- Entropy
- Softmax
- Cross-Entropy Loss
- Logistic Regression
- The Perceptron
- Activation Functions
- The Artificial Neuron
- Matrix Multiplication
- Multilayer Perceptron (MLP)
- The Chain Rule
- The Forward Pass
- Computational Graphs and Autodiff
- Backpropagation
- Convolutional Neural Networks
- Embeddings
- Attention
- Self-Attention
- Vision Transformer (ViT)
- Audio and Spectrograms
- Turning Signals into Tokens
A sequence of pieces
Every model since Chapter 7 reads a sequence of vectors. Text was easy to cut up: tokenization splits it into subwords. For other signals the recipe is the same idea with a different knife:
| Signal | Cut into | One token is |
|---|---|---|
| Image | p × p patches | one patch |
| Audio | short frames (via a spectrogram or a learned codec) | about 10–80 ms of sound |
| Video | spacetime patches | a patch across a few frames |
| Robot action | one number per joint or axis, binned | one bin index |
The Vision Transformer made the image case concrete: a 224 × 224 image in 16 × 16 patches is 14 × 14 = 196 tokens, each a flattened 16 × 16 × 3 = 768-number patch multiplied by a learned matrix.
Continuous or discrete
There are two ways to turn a piece into a token.
Continuous. Multiply the piece's numbers by a learned matrix (or run it through an encoder) and pass the vector on. Nothing is rounded off, so the model keeps whatever detail the encoder kept. This is what models that only need to read images do: ViT, CLIP's image encoder, and the image side of most vision-language models.
Discrete. Learn a codebook, a list of K reference vectors, and replace each piece with the index of its nearest entry. The image becomes a list of integers, just like text. VQ-VAE introduced this learned discrete bottleneck for images, video and speech. Established The payoff is generation: a model that predicts the next token can now predict the next image token. DALL·E compressed each 256 × 256 image into a 32 × 32 grid of tokens, each one of 8,192 possible values, and trained a Transformer on up to 256 text tokens followed by those 1,024 image tokens. Established
Discrete tokens are lossy compression. A codebook entry is an average of many patches, so fine detail (thin strokes, small text) is rounded away first.
Tiny example (lab numbers)
The lab uses 64 × 64 pictures:
- 16-pixel patches: 4 × 4 = 16 tokens, and 16² = 256 attention pairs.
- 8-pixel patches: 64 tokens, 4,096 pairs.
- 4-pixel patches: 256 tokens, 65,536 pairs.
As codes, 16 tokens from a 128-entry codebook are 16 × log₂128 = 112 bits, against 98,304 bits of raw 8-bit pixels. That is enormous compression, and the reconstruction shows its cost on the lettering. With 16 tokens and 8 codes, the chart's letters are off by 35% on average (as a share of full brightness), against 12% for the rest of the picture. With 256 tokens and 128 codes the rest is almost exact (0.4%), but the letters are still off by 11.5%. In every setting the lab offers, the error on the text is higher than on the rest of the picture.
The cost is the token count
Tokens from images are billed, cached and attended to like text tokens. Halving the patch size quadruples the token count, and full attention compares every pair, so its cost grows with the square (attention complexity). The KV cache grows linearly with them too.
This is why high-resolution images and long videos are expensive for multimodal models, and why systems shrink, crop, tile or compress visual input before the language model sees it. Interpretation A data-engineering analogy: choosing the patch size is choosing a partition size. Smaller partitions keep more detail per row but multiply the rows every join has to touch.
Beyond pictures
- Audio usually starts as a spectrogram (Chapter 5) or the codes of a neural audio codec. Whisper's encoder, for example, sees 30-second chunks as 80-channel log-mel spectrograms with a 10 ms stride.
- Video adds time: a patch spans a small block of pixels across a few frames. OpenAI's Sora report describes compressing video into a latent space and cutting it into spacetime patches that act as Transformer tokens. Established
- Actions are numbers too. RT-2 discretized each continuous robot-action dimension into 256 bins and wrote an action as 8 integers. Established
Mini experiment
In the lab, choose the receipt and 8-pixel patches. Compare 8 and 128 codes, then compare 16-pixel and 4-pixel patches at 32 codes. Which change helps the text more per extra bit? Then switch to raw patches and explain why their error is always zero, and what they cost instead.
Why should I care?
As a researcher
How a modality is tokenized decides what the model can perceive (a 16-pixel patch hides 3-pixel text), how long its sequences are, and whether it can generate the modality at all.
As an engineer
Image and audio tokens are billed and cached like text. A screenshot can cost hundreds or thousands of tokens, which sets your latency, context budget and price.
Modern systems that depend on it
- Vision Transformers and CLIP
- vision-language models
- token-based image and audio generation
- diffusion transformers
- vision-language-action models
Historical context
Before
Each modality had its own architecture: CNNs for images, recurrent or convolutional models for audio, separate pipelines that could not share a model with text.
After
Images, audio, video and robot actions all arrive as token sequences, so one Transformer architecture, and sometimes one model, can process them together.
Used today
Vision-language models read images as patch tokens, voice models read and write audio codec tokens, and video generators denoise spacetime patches.
What to remember
- Image: p×p patches; a 224×224 image with 16-pixel patches gives 14 × 14 = 196 tokens.
- Continuous tokens keep each patch's information as a vector; discrete tokens replace it with a codebook index.
- Discrete tokens let a model generate a modality exactly as it generates words (DALL·E: 32 × 32 = 1,024 image tokens from 8,192 codes).
- Halving the patch size quadruples the tokens; attention cost grows with the square of the token count.
- Audio: frames from a spectrogram or a neural codec; video: patches through space and time.
- Small details (text, thin lines) are the first casualty of coarse patches or small codebooks.
Key papers
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Alexey Dosovitskiy, Lucas Beyer et al. · 2020 · ICLR 2021
Showed that a nearly unmodified Transformer, reading an image as a sequence of patches, can match strong CNNs when pretrained on enough data. Vision and language began to share one architecture.
How to read it: The key result is the comparison across pretraining dataset sizes: with less data CNNs win, with more the ViT catches up.
Neural Discrete Representation Learning
Aaron van den Oord, Oriol Vinyals, Koray Kavukcuoglu · 2017
VQ-VAE turned images, video and speech into sequences of discrete codes from a learned codebook. That is the trick that lets a language-model-style Transformer generate pictures and sound as tokens.
How to read it: Focus on the nearest-neighbour lookup and the straight-through gradient: the encoder never sees the rounding, it receives the decoder's gradient as if no rounding had happened.
Zero-Shot Text-to-Image Generation
Aditya Ramesh, Mikhail Pavlov et al. · 2021
DALL·E treated text-to-image generation as language modelling over one stream of text tokens followed by image tokens, and showed that scale made it work.
How to read it: Section 2 is the recipe in two stages. Count the tokens: 256 + 1,024 per example is why image tokens are expensive context.
SoundStream: An End-to-End Neural Audio Codec
Neil Zeghidour, Alejandro Luebs et al. · 2021
A neural codec that turns audio into a short stream of discrete codes. Codecs like it are how audio language models read and write sound as tokens.
How to read it: Residual vector quantization is the idea to take away: quantize, subtract, quantize the remainder again. Each extra codebook adds detail and bitrate.
A Generalist Agent
Scott Reed, Konrad Zolna et al. · 2022
Gato put text, images, game buttons and robot joint torques into one token sequence for one network, an early test of the 'everything is tokens' idea for action.
How to read it: Read the tokenization section (how continuous values become tokens) and compare per-task results with specialists before reading 'generalist' as 'good at everything'.
Chameleon: Mixed-Modal Early-Fusion Foundation Models
Chameleon Team · 2024
A clear example of early fusion: images become discrete tokens in the same sequence and vocabulary as text, so one model reads and writes both.
How to read it: The stability section is the practical lesson: mixing modalities in one sequence caused training divergences that needed architectural fixes.