Skip to content
Road to Intelligence

Concept · Chapter 15: Multimodal AI

Vision-Language-Action Models

Should knowUnderstand11 minDifficulty

A vision-language-action model is a vision-language model fine-tuned to output robot actions, often written as discretized tokens, from camera images and a natural-language instruction.

The problem

Robot policies trained only on robot data, which is scarce and expensive, generalize poorly to new objects, scenes and instructions.

The solution

Start from a vision-language model pretrained on web images and text, represent each action as a few tokens (one bin per action dimension), and fine-tune on robot demonstrations, often mixed with web data.

The consequence

Some web knowledge transfers into control (new objects, simple semantic reasoning), but real-world success rates, latency and safety are still far from the reliability of text models.

Actions are just more tokens

A robot arm's action at one moment is a short list of numbers: how far to move and rotate the gripper, how open to make it, whether to stop. RT-2 represented the 6-DoF position and rotation change of the end-effector, the gripper extension and a terminate command as 8 integers, discretizing each continuous dimension into 256 uniform bins, and mapped those integers onto tokens in the vision-language model's existing vocabulary. Established The model reads a camera image and an instruction and "writes" the next action, then the next, in a loop.

Gato had earlier put text, images, Atari button presses and robot joint torques into one token sequence for a single network with one set of weights. Established

Why start from a VLM

Robot demonstrations are scarce; image–text pairs are not. RT-2 co-fine-tuned large vision-language models on robot trajectories and on web vision-language tasks, and in about 6,000 evaluation trials showed improved generalization to novel objects and the ability to follow commands not present in the robot data, such as placing an object on a particular number or icon. Established What transfers is mostly recognition and semantics (what "the smallest object" or "an improvised hammer" refers to), not new motor skill: the motions themselves still come from robot demonstrations. Interpretation

OpenVLA, a 7B-parameter open model built on Llama 2 with a visual encoder fusing DINOv2 and SigLIP features, was trained on 970k real-world robot episodes and reported 16.5% higher absolute task success than the 55B RT-2-X across 29 tasks. Established It can be fine-tuned with LoRA and served quantized, the same tools as Chapter 11.

Tiny example

With 256 bins per dimension, a gripper that can move 20 cm along one axis has a resolution of 20/256 ≈ 0.8 mm per bin. Fine for picking up a can; coarse for threading a needle. Choosing bin ranges and counts is a design decision in the same way that choosing a patch size is.

What is hard

  • Data: each demonstration needs a robot, a person and time; datasets are orders of magnitude smaller than web text.
  • Closed loop: a wrong action changes the world, and the next image shows the consequences. Errors compound, as in agents.
  • Latency and safety: a large model running at a few actions per second limits fast motions; mistakes have physical consequences.
  • Evaluation: success rates come from a modest number of real trials per task, and results depend on the robot, scene and task selection.

Mini experiment

Write down the 8 numbers RT-2's format would need for "move the gripper 2 cm to the left and close it". Which of them are continuous, which are discrete, and what does the 256-bin limit mean for a task that needs 0.1 mm precision?

What to remember

  • Action as text: each continuous action dimension is cut into 256 bins; RT-2 writes an action as 8 integers.
  • Co-fine-tuning on robot data plus web vision-language data keeps the web knowledge.
  • RT-2 was evaluated in about 6,000 real trials; OpenVLA (7B, open) trained on 970k robot episodes.
  • Generalization is reported per category (new objects, backgrounds, instructions), not as one number.
  • Robot data is the bottleneck: the web has text and images, not joint torques.

Key papers

Optional

A Generalist Agent

Scott Reed, Konrad Zolna et al. · 2022

Gato put text, images, game buttons and robot joint torques into one token sequence for one network, an early test of the 'everything is tokens' idea for action.

How to read it: Read the tokenization section (how continuous values become tokens) and compare per-task results with specialists before reading 'generalist' as 'good at everything'.

~40 min readarXiv:2205.06175✓ verified 2026-10-07
Essential

RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Anthony Brohan, Noah Brown et al. · 2023

RT-2 named and demonstrated the vision-language-action model: a VLM fine-tuned to write robot actions as tokens, carrying some web knowledge into control.

How to read it: Section 3.2 is the action-as-text trick. Then read the generalization results with their categories (unseen objects, backgrounds, environments) rather than one average.

~45 min readarXiv:2307.15818✓ verified 2026-10-07
Optional

OpenVLA: An Open-Source Vision-Language-Action Model

Moo Jin Kim, Karl Pertsch et al. · 2024

An open 7B vision-language-action model, so the RT-2 recipe could be studied, fine-tuned and served outside one lab.

How to read it: Note how evaluation works in robotics: a fixed set of tasks on real robots, success rates from a modest number of trials each.

~40 min readarXiv:2406.09246✓ verified 2026-10-07