Skip to content
Road to Intelligence

Concept · Chapter 15: Multimodal AI

World Models

FrontierUnderstand11 minDifficulty

A world model is a learned model that predicts how an environment will change in response to actions, so that an agent can plan or practise inside its predictions instead of only in the real world.

The problem

Learning to act by trial and error in the real world is slow, expensive and sometimes dangerous; every attempt costs a real interaction.

The solution

Learn a compressed model of the environment's dynamics from experience or video, then imagine rollouts inside it to train a policy or to compare plans before acting.

The consequence

Agents can learn with far fewer real interactions and can be trained on imagined experience, but they exploit the model's mistakes, and whether large video generators amount to world models is an open question.

Imagine before you act

Reinforcement learning in Chapter 5 learned from real consequences. A world model adds a step: learn how the environment behaves, then use that model to imagine. Ha and Schmidhuber trained a VAE to compress game frames and a recurrent network to predict the next compressed frame, then trained a very small controller on these features, even entirely inside the model's own generated "dream", and transferred it back to the real environment. Established

They also observed agents exploiting imperfections of the world model, and found that sampling the model at a higher temperature (more uncertainty) made it harder to exploit and gave policies that transferred better. Established This is Goodhart's law again: optimize hard against a learned stand-in and you find its flaws.

From games to general methods

DreamerV3 learns a world model and improves its behaviour by imagining future scenarios; with a single configuration it outperformed specialized methods across over 150 diverse tasks and was the first algorithm to collect diamonds in Minecraft from scratch without human data or curricula. Established Genie, an 11B-parameter model trained on 30,000 hours of internet gameplay video from hundreds of 2D platformer games, generates environments a user can act in frame by frame, learning a small latent action space without any action labels. Established

Are video generators world models?

A video generator predicts plausible next frames; a world model must predict the consequences of actions, accurately enough to plan with. OpenAI's Sora report argues that scaling video models is a promising path toward simulators of the physical world. Speculative The same report lists failures with basic physics. Plausible-looking frames are a weaker test than correct predictions under intervention; to evaluate a world model, change the action and check whether the predicted outcome changes the way the real world would. Interpretation

Tiny example

An agent plans 5 steps ahead with 4 possible actions each: 4⁵ = 1,024 imagined rollouts per decision. Inside a learned model each rollout costs a few network calls and no real-world risk. In the real world it would cost 1,024 trial runs. The saving is real only while the model's predictions stay accurate over 5 steps; small errors compound with each imagined step.

Mini experiment

Take the gridworld from Chapter 5. Suppose you learned its transitions from 20 random episodes and then trained Q-learning only inside that learned model. Which squares would the model get wrong (never visited?), and how could the agent exploit those mistakes?

What to remember

  • World model = learned p(next state | state, action); the agent can plan or train inside it.
  • Ha and Schmidhuber: VAE for frames, RNN for dynamics, a tiny controller trained in the 'dream'.
  • Agents exploit flaws in their own world model; adding uncertainty (temperature) makes the dream harder to cheat.
  • DreamerV3: one configuration across 150+ tasks; diamonds in Minecraft from scratch.
  • Genie learned controllable environments from unlabelled video by inferring latent actions.

Key papers

Important

World Models

David Ha, Jürgen Schmidhuber · 2018

A clear, small demonstration of the world-model idea: learn a compressed model of an environment, then train a tiny controller inside the model's own imagined rollouts.

How to read it: The interactive version at worldmodels.github.io is the best way in. Watch for where the controller exploits flaws in its own dream, and how the authors counter it with extra randomness.

~25 min readarXiv:1803.10122✓ verified 2026-10-07
Optional

Mastering Diverse Domains through World Models

Danijar Hafner, Jurgis Pasukonis et al. · 2023

DreamerV3 made learning inside a world model a general-purpose RL method: one configuration across more than 150 tasks.

How to read it: Read the overview figure, then the robustness techniques: the practical difficulty of world models is keeping learning stable across very different reward scales.

~50 min readarXiv:2301.04104✓ verified 2026-10-07
Optional

Genie: Generative Interactive Environments

Jake Bruce, Michael Dennis et al. · 2024

A world model learned from unlabelled internet video that you can act in frame by frame, a step from 'generate a video' to 'generate an environment'.

How to read it: The latent action model is the novel part: it infers a small set of discrete 'actions' from what changes between frames.

~45 min readarXiv:2402.15391✓ verified 2026-10-07