Skip to content
Road to Intelligence

Concept · Chapter 15: Multimodal AI

Turning Signals into Tokens

Must knowKnow well14 minDifficulty

Before a Transformer can use an image, a sound or a video, the signal is cut into pieces (patches, frames, spacetime blocks) and each piece becomes one token: either a continuous vector or the index of an entry in a learned codebook.

The problem

A Transformer reads a sequence of vectors, but a photo is a grid of millions of numbers, audio is tens of thousands of samples per second, and video is both at once.

The solution

Cut the signal into regular pieces, embed each piece as one vector (continuous tokens), or snap it to the nearest entry of a learned codebook and use the entry's number (discrete tokens).

The consequence

Every modality becomes a sequence, so the machinery of Chapters 7–11 applies. The price is token count: images and video are expensive context, and discrete codes lose detail.

A sequence of pieces

Every model since Chapter 7 reads a sequence of vectors. Text was easy to cut up: tokenization splits it into subwords. For other signals the recipe is the same idea with a different knife:

SignalCut intoOne token is
Imagep × p patchesone patch
Audioshort frames (via a spectrogram or a learned codec)about 10–80 ms of sound
Videospacetime patchesa patch across a few frames
Robot actionone number per joint or axis, binnedone bin index

The Vision Transformer made the image case concrete: a 224 × 224 image in 16 × 16 patches is 14 × 14 = 196 tokens, each a flattened 16 × 16 × 3 = 768-number patch multiplied by a learned matrix.

Continuous or discrete

There are two ways to turn a piece into a token.

Continuous. Multiply the piece's numbers by a learned matrix (or run it through an encoder) and pass the vector on. Nothing is rounded off, so the model keeps whatever detail the encoder kept. This is what models that only need to read images do: ViT, CLIP's image encoder, and the image side of most vision-language models.

Discrete. Learn a codebook, a list of K reference vectors, and replace each piece with the index of its nearest entry. The image becomes a list of integers, just like text. VQ-VAE introduced this learned discrete bottleneck for images, video and speech. Established The payoff is generation: a model that predicts the next token can now predict the next image token. DALL·E compressed each 256 × 256 image into a 32 × 32 grid of tokens, each one of 8,192 possible values, and trained a Transformer on up to 256 text tokens followed by those 1,024 image tokens. Established

Discrete tokens are lossy compression. A codebook entry is an average of many patches, so fine detail (thin strokes, small text) is rounded away first.

Tiny example (lab numbers)

The lab uses 64 × 64 pictures:

  • 16-pixel patches: 4 × 4 = 16 tokens, and 16² = 256 attention pairs.
  • 8-pixel patches: 64 tokens, 4,096 pairs.
  • 4-pixel patches: 256 tokens, 65,536 pairs.

As codes, 16 tokens from a 128-entry codebook are 16 × log₂128 = 112 bits, against 98,304 bits of raw 8-bit pixels. That is enormous compression, and the reconstruction shows its cost on the lettering. With 16 tokens and 8 codes, the chart's letters are off by 35% on average (as a share of full brightness), against 12% for the rest of the picture. With 256 tokens and 128 codes the rest is almost exact (0.4%), but the letters are still off by 11.5%. In every setting the lab offers, the error on the text is higher than on the rest of the picture.

The cost is the token count

Tokens from images are billed, cached and attended to like text tokens. Halving the patch size quadruples the token count, and full attention compares every pair, so its cost grows with the square (attention complexity). The KV cache grows linearly with them too.

This is why high-resolution images and long videos are expensive for multimodal models, and why systems shrink, crop, tile or compress visual input before the language model sees it. Interpretation A data-engineering analogy: choosing the patch size is choosing a partition size. Smaller partitions keep more detail per row but multiply the rows every join has to touch.

Beyond pictures

  • Audio usually starts as a spectrogram (Chapter 5) or the codes of a neural audio codec. Whisper's encoder, for example, sees 30-second chunks as 80-channel log-mel spectrograms with a 10 ms stride.
  • Video adds time: a patch spans a small block of pixels across a few frames. OpenAI's Sora report describes compressing video into a latent space and cutting it into spacetime patches that act as Transformer tokens. Established
  • Actions are numbers too. RT-2 discretized each continuous robot-action dimension into 256 bins and wrote an action as 8 integers. Established

Mini experiment

In the lab, choose the receipt and 8-pixel patches. Compare 8 and 128 codes, then compare 16-pixel and 4-pixel patches at 32 codes. Which change helps the text more per extra bit? Then switch to raw patches and explain why their error is always zero, and what they cost instead.

Why should I care?

As a researcher

How a modality is tokenized decides what the model can perceive (a 16-pixel patch hides 3-pixel text), how long its sequences are, and whether it can generate the modality at all.

As an engineer

Image and audio tokens are billed and cached like text. A screenshot can cost hundreds or thousands of tokens, which sets your latency, context budget and price.

Modern systems that depend on it

  • Vision Transformers and CLIP
  • vision-language models
  • token-based image and audio generation
  • diffusion transformers
  • vision-language-action models

Historical context

Before

Each modality had its own architecture: CNNs for images, recurrent or convolutional models for audio, separate pipelines that could not share a model with text.

After

Images, audio, video and robot actions all arrive as token sequences, so one Transformer architecture, and sometimes one model, can process them together.

Used today

Vision-language models read images as patch tokens, voice models read and write audio codec tokens, and video generators denoise spacetime patches.

What to remember

  • Image: p×p patches; a 224×224 image with 16-pixel patches gives 14 × 14 = 196 tokens.
  • Continuous tokens keep each patch's information as a vector; discrete tokens replace it with a codebook index.
  • Discrete tokens let a model generate a modality exactly as it generates words (DALL·E: 32 × 32 = 1,024 image tokens from 8,192 codes).
  • Halving the patch size quadruples the tokens; attention cost grows with the square of the token count.
  • Audio: frames from a spectrogram or a neural codec; video: patches through space and time.
  • Small details (text, thin lines) are the first casualty of coarse patches or small codebooks.

Key papers

Essential

An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Alexey Dosovitskiy, Lucas Beyer et al. · 2020 · ICLR 2021

Showed that a nearly unmodified Transformer, reading an image as a sequence of patches, can match strong CNNs when pretrained on enough data. Vision and language began to share one architecture.

How to read it: The key result is the comparison across pretraining dataset sizes: with less data CNNs win, with more the ViT catches up.

~40 min readarXiv:2010.11929✓ verified 2026-09-26
Important

Neural Discrete Representation Learning

Aaron van den Oord, Oriol Vinyals, Koray Kavukcuoglu · 2017

VQ-VAE turned images, video and speech into sequences of discrete codes from a learned codebook. That is the trick that lets a language-model-style Transformer generate pictures and sound as tokens.

How to read it: Focus on the nearest-neighbour lookup and the straight-through gradient: the encoder never sees the rounding, it receives the decoder's gradient as if no rounding had happened.

~30 min readarXiv:1711.00937✓ verified 2026-10-07
Important

Zero-Shot Text-to-Image Generation

Aditya Ramesh, Mikhail Pavlov et al. · 2021

DALL·E treated text-to-image generation as language modelling over one stream of text tokens followed by image tokens, and showed that scale made it work.

How to read it: Section 2 is the recipe in two stages. Count the tokens: 256 + 1,024 per example is why image tokens are expensive context.

~40 min readarXiv:2102.12092✓ verified 2026-10-07
Optional

SoundStream: An End-to-End Neural Audio Codec

Neil Zeghidour, Alejandro Luebs et al. · 2021

A neural codec that turns audio into a short stream of discrete codes. Codecs like it are how audio language models read and write sound as tokens.

How to read it: Residual vector quantization is the idea to take away: quantize, subtract, quantize the remainder again. Each extra codebook adds detail and bitrate.

~35 min readarXiv:2107.03312✓ verified 2026-10-07
Optional

A Generalist Agent

Scott Reed, Konrad Zolna et al. · 2022

Gato put text, images, game buttons and robot joint torques into one token sequence for one network, an early test of the 'everything is tokens' idea for action.

How to read it: Read the tokenization section (how continuous values become tokens) and compare per-task results with specialists before reading 'generalist' as 'good at everything'.

~40 min readarXiv:2205.06175✓ verified 2026-10-07
Optional

Chameleon: Mixed-Modal Early-Fusion Foundation Models

Chameleon Team · 2024

A clear example of early fusion: images become discrete tokens in the same sequence and vocabulary as text, so one model reads and writes both.

How to read it: The stability section is the practical lesson: mixing modalities in one sequence caused training divergences that needed architectural fixes.

~45 min readarXiv:2405.09818✓ verified 2026-10-07