Skip to content
Road to Intelligence

Concept · Chapter 15: Multimodal AI

Video Generation

Should knowUnderstand10 minDifficulty

Video generators extend image diffusion through time: they compress video into a latent, cut it into spacetime patches, and denoise all frames together so motion stays consistent.

The problem

Generating frames one at a time with an image model gives flicker and drift: nothing ties frame 40 to frame 1.

The solution

Treat a clip as one object. Compress it in space and time, denoise the whole spacetime latent jointly with a Transformer over spacetime patches, and condition on text, images or earlier frames.

The consequence

Short, coherent, high-resolution clips became possible, at very large compute cost per second of video, with persistent failures in physics, object permanence and long-range consistency.

Time is one more axis

An image latent is a grid of positions; a video latent adds time. Ho and colleagues extended the standard image diffusion architecture to video and found that training jointly on images and video helped optimization. Established Denoising every frame together, with attention across time, is what makes a ball stay one ball as it moves.

OpenAI's February 2024 Sora report describes compressing video into a lower-dimensional latent in both space and time, cutting it into spacetime patches that act as Transformer tokens, and training a diffusion transformer on videos and images of variable duration, resolution and aspect ratio. Established Because an image is a video with one frame, one model learns from both. The report gives no model or implementation details.

Tiny example: the token bill

Suppose a compressor reduces each side by 8 and time by 4, and the Transformer uses 2 × 2 patches in the latent. A 5-second clip at 24 frames per second and 512 × 512 pixels has 120 frames → 30 latent frames of 64 × 64 → 30 × 32 × 32 = 30,720 tokens. Double the resolution and the count quadruples; double the duration and it doubles. With full attention, cost grows with the square. These factors are illustrative, not any specific system's, but the arithmetic is why video models compress aggressively and why minute-long clips are expensive.

What still goes wrong

The Sora report itself says the model does not accurately model the physics of many basic interactions, such as glass shattering, does not always produce correct changes of object state, and develops incoherence in long samples, with objects appearing spontaneously. Established The same report suggests that scaling video models is a promising path toward general-purpose simulators of the physical world; whether video prediction alone yields reliable physical understanding is an open question. Speculative

Mini experiment

Using the token bill above, compute the tokens for a 60-second clip at the same settings, and the number of attention pairs. Then halve the time compression (4 → 2): which costs more, doubling duration or halving temporal compression?

What to remember

  • A video is a 3-D grid (time × height × width); patches can span a few frames as well as pixels.
  • Denoising all frames together is what keeps motion coherent.
  • Training on images and videos together helps; an image is a one-frame video.
  • Token counts explode with duration and resolution, so compression in time matters as much as in space.
  • Plausible-looking motion is not correct physics: Sora's own report lists failures such as glass shattering.

Key papers

Optional

Video Diffusion Models

Jonathan Ho, Tim Salimans et al. · 2022

An early, clean extension of image diffusion to video, including joint training on images and videos.

How to read it: Notice how much is reused from images: the step from image to video here is mostly an architecture change, not a new generative principle.

~30 min readarXiv:2204.03458✓ verified 2026-10-07
Important

Scalable Diffusion Models with Transformers

William Peebles, Saining Xie · 2022

The Diffusion Transformer (DiT) replaced the U-Net with a Transformer over latent patches; OpenAI's Sora report describes Sora as a diffusion transformer.

How to read it: The Gflops-versus-FID plot is the argument: the familiar language-model scaling story, now for a denoiser.

~35 min readarXiv:2212.09748✓ verified 2026-10-07
Optional

Genie: Generative Interactive Environments

Jake Bruce, Michael Dennis et al. · 2024

A world model learned from unlabelled internet video that you can act in frame by frame, a step from 'generate a video' to 'generate an environment'.

How to read it: The latent action model is the novel part: it infers a small set of discrete 'actions' from what changes between frames.

~45 min readarXiv:2402.15391✓ verified 2026-10-07