Concept · Chapter 15: Multimodal AI
Video Generation
Video generators extend image diffusion through time: they compress video into a latent, cut it into spacetime patches, and denoise all frames together so motion stays consistent.
The problem
Generating frames one at a time with an image model gives flicker and drift: nothing ties frame 40 to frame 1.
The solution
Treat a clip as one object. Compress it in space and time, denoise the whole spacetime latent jointly with a Transformer over spacetime patches, and condition on text, images or earlier frames.
The consequence
Short, coherent, high-resolution clips became possible, at very large compute cost per second of video, with persistent failures in physics, object permanence and long-range consistency.
You should understand first
- Probability and Distributions
- Text as Data
- Vectors
- One-Hot Encoding
- Tokenization
- Tensors and Shapes
- Images as Tensors
- Dot Product
- Convolution
- The Turing Test
- Symbolic AI
- Logic and Rules
- Expert Systems
- Knowledge Representation
- The Knowledge-Acquisition Bottleneck
- From Rules to Learning
- Supervised, Unsupervised and Self-Supervised Learning
- Features, Labels and Tasks
- Loss Functions
- Derivatives and Gradients
- Gradient Descent
- Linear Regression
- Entropy
- Softmax
- Cross-Entropy Loss
- Logistic Regression
- The Perceptron
- Activation Functions
- The Artificial Neuron
- Matrix Multiplication
- Multilayer Perceptron (MLP)
- The Chain Rule
- The Forward Pass
- Computational Graphs and Autodiff
- Backpropagation
- Convolutional Neural Networks
- Embeddings
- Attention
- Self-Attention
- Vision Transformer (ViT)
- Audio and Spectrograms
- Turning Signals into Tokens
- Generative Models: GANs, VAEs and Image Tokens
- Diffusion Models
- Classifier-Free Guidance
- Multi-Head Attention
- Causal Masking
- Positional Encoding
- Residual Connections
- Layer Normalization
- Feed-Forward Sublayer (MLP)
- The Transformer Block
- Encoder, Decoder & Encoder–Decoder
- Text Embeddings
- Contrastive Learning
- CLIP and Image–Text Embeddings
- Latent Diffusion and Text-to-Image
- Video Generation
Time is one more axis
An image latent is a grid of positions; a video latent adds time. Ho and colleagues extended the standard image diffusion architecture to video and found that training jointly on images and video helped optimization. Established Denoising every frame together, with attention across time, is what makes a ball stay one ball as it moves.
OpenAI's February 2024 Sora report describes compressing video into a lower-dimensional latent in both space and time, cutting it into spacetime patches that act as Transformer tokens, and training a diffusion transformer on videos and images of variable duration, resolution and aspect ratio. Established Because an image is a video with one frame, one model learns from both. The report gives no model or implementation details.
Tiny example: the token bill
Suppose a compressor reduces each side by 8 and time by 4, and the Transformer uses 2 × 2 patches in the latent. A 5-second clip at 24 frames per second and 512 × 512 pixels has 120 frames → 30 latent frames of 64 × 64 → 30 × 32 × 32 = 30,720 tokens. Double the resolution and the count quadruples; double the duration and it doubles. With full attention, cost grows with the square. These factors are illustrative, not any specific system's, but the arithmetic is why video models compress aggressively and why minute-long clips are expensive.
What still goes wrong
The Sora report itself says the model does not accurately model the physics of many basic interactions, such as glass shattering, does not always produce correct changes of object state, and develops incoherence in long samples, with objects appearing spontaneously. Established The same report suggests that scaling video models is a promising path toward general-purpose simulators of the physical world; whether video prediction alone yields reliable physical understanding is an open question. SpeculativeMini experiment
Using the token bill above, compute the tokens for a 60-second clip at the same settings, and the number of attention pairs. Then halve the time compression (4 → 2): which costs more, doubling duration or halving temporal compression?
What to remember
- A video is a 3-D grid (time × height × width); patches can span a few frames as well as pixels.
- Denoising all frames together is what keeps motion coherent.
- Training on images and videos together helps; an image is a one-frame video.
- Token counts explode with duration and resolution, so compression in time matters as much as in space.
- Plausible-looking motion is not correct physics: Sora's own report lists failures such as glass shattering.
Key papers
Video Diffusion Models
Jonathan Ho, Tim Salimans et al. · 2022
An early, clean extension of image diffusion to video, including joint training on images and videos.
How to read it: Notice how much is reused from images: the step from image to video here is mostly an architecture change, not a new generative principle.
Scalable Diffusion Models with Transformers
William Peebles, Saining Xie · 2022
The Diffusion Transformer (DiT) replaced the U-Net with a Transformer over latent patches; OpenAI's Sora report describes Sora as a diffusion transformer.
How to read it: The Gflops-versus-FID plot is the argument: the familiar language-model scaling story, now for a denoiser.
Genie: Generative Interactive Environments
Jake Bruce, Michael Dennis et al. · 2024
A world model learned from unlabelled internet video that you can act in frame by frame, a step from 'generate a video' to 'generate an environment'.
How to read it: The latent action model is the novel part: it infers a small set of discrete 'actions' from what changes between frames.