Concept · Chapter 15: Multimodal AI
Vision-Language-Action Models
A vision-language-action model is a vision-language model fine-tuned to output robot actions, often written as discretized tokens, from camera images and a natural-language instruction.
The problem
Robot policies trained only on robot data, which is scarce and expensive, generalize poorly to new objects, scenes and instructions.
The solution
Start from a vision-language model pretrained on web images and text, represent each action as a few tokens (one bin per action dimension), and fine-tune on robot demonstrations, often mixed with web data.
The consequence
Some web knowledge transfers into control (new objects, simple semantic reasoning), but real-world success rates, latency and safety are still far from the reliability of text models.
You should understand first
- Vectors
- Dot Product
- Embeddings
- Attention
- Probability and Distributions
- Softmax
- Self-Attention
- Multi-Head Attention
- Causal Masking
- Positional Encoding
- Residual Connections
- Layer Normalization
- Feed-Forward Sublayer (MLP)
- The Transformer Block
- Encoder, Decoder & Encoder–Decoder
- Text Embeddings
- Entropy
- Loss Functions
- Cross-Entropy Loss
- Contrastive Learning
- Text as Data
- One-Hot Encoding
- Tokenization
- Tensors and Shapes
- Images as Tensors
- Convolution
- The Turing Test
- Symbolic AI
- Logic and Rules
- Expert Systems
- Knowledge Representation
- The Knowledge-Acquisition Bottleneck
- From Rules to Learning
- Supervised, Unsupervised and Self-Supervised Learning
- Features, Labels and Tasks
- Derivatives and Gradients
- Gradient Descent
- Linear Regression
- Logistic Regression
- The Perceptron
- Activation Functions
- The Artificial Neuron
- Matrix Multiplication
- Multilayer Perceptron (MLP)
- The Chain Rule
- The Forward Pass
- Computational Graphs and Autodiff
- Backpropagation
- Convolutional Neural Networks
- Vision Transformer (ViT)
- Audio and Spectrograms
- Turning Signals into Tokens
- CLIP and Image–Text Embeddings
- Conditional Probability and Bayes' Theorem
- Probability of Sequences
- Language Modeling
- Pretrain, Then Fine-Tune
- Supervised Fine-Tuning
- Vision-Language Models
- Expected Value and Variance
- Reinforcement Learning
- Vision-Language-Action Models
Actions are just more tokens
A robot arm's action at one moment is a short list of numbers: how far to move and rotate the gripper, how open to make it, whether to stop. RT-2 represented the 6-DoF position and rotation change of the end-effector, the gripper extension and a terminate command as 8 integers, discretizing each continuous dimension into 256 uniform bins, and mapped those integers onto tokens in the vision-language model's existing vocabulary. Established The model reads a camera image and an instruction and "writes" the next action, then the next, in a loop.
Gato had earlier put text, images, Atari button presses and robot joint torques into one token sequence for a single network with one set of weights. EstablishedWhy start from a VLM
Robot demonstrations are scarce; image–text pairs are not. RT-2 co-fine-tuned large vision-language models on robot trajectories and on web vision-language tasks, and in about 6,000 evaluation trials showed improved generalization to novel objects and the ability to follow commands not present in the robot data, such as placing an object on a particular number or icon. Established What transfers is mostly recognition and semantics (what "the smallest object" or "an improvised hammer" refers to), not new motor skill: the motions themselves still come from robot demonstrations. Interpretation
OpenVLA, a 7B-parameter open model built on Llama 2 with a visual encoder fusing DINOv2 and SigLIP features, was trained on 970k real-world robot episodes and reported 16.5% higher absolute task success than the 55B RT-2-X across 29 tasks. Established It can be fine-tuned with LoRA and served quantized, the same tools as Chapter 11.
Tiny example
With 256 bins per dimension, a gripper that can move 20 cm along one axis has a resolution of 20/256 ≈ 0.8 mm per bin. Fine for picking up a can; coarse for threading a needle. Choosing bin ranges and counts is a design decision in the same way that choosing a patch size is.
What is hard
- Data: each demonstration needs a robot, a person and time; datasets are orders of magnitude smaller than web text.
- Closed loop: a wrong action changes the world, and the next image shows the consequences. Errors compound, as in agents.
- Latency and safety: a large model running at a few actions per second limits fast motions; mistakes have physical consequences.
- Evaluation: success rates come from a modest number of real trials per task, and results depend on the robot, scene and task selection.
Mini experiment
Write down the 8 numbers RT-2's format would need for "move the gripper 2 cm to the left and close it". Which of them are continuous, which are discrete, and what does the 256-bin limit mean for a task that needs 0.1 mm precision?
What to remember
- Action as text: each continuous action dimension is cut into 256 bins; RT-2 writes an action as 8 integers.
- Co-fine-tuning on robot data plus web vision-language data keeps the web knowledge.
- RT-2 was evaluated in about 6,000 real trials; OpenVLA (7B, open) trained on 970k robot episodes.
- Generalization is reported per category (new objects, backgrounds, instructions), not as one number.
- Robot data is the bottleneck: the web has text and images, not joint torques.
Key papers
A Generalist Agent
Scott Reed, Konrad Zolna et al. · 2022
Gato put text, images, game buttons and robot joint torques into one token sequence for one network, an early test of the 'everything is tokens' idea for action.
How to read it: Read the tokenization section (how continuous values become tokens) and compare per-task results with specialists before reading 'generalist' as 'good at everything'.
RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
Anthony Brohan, Noah Brown et al. · 2023
RT-2 named and demonstrated the vision-language-action model: a VLM fine-tuned to write robot actions as tokens, carrying some web knowledge into control.
How to read it: Section 3.2 is the action-as-text trick. Then read the generalization results with their categories (unseen objects, backgrounds, environments) rather than one average.
OpenVLA: An Open-Source Vision-Language-Action Model
Moo Jin Kim, Karl Pertsch et al. · 2024
An open 7B vision-language-action model, so the RT-2 recipe could be studied, fine-tuned and served outside one lab.
How to read it: Note how evaluation works in robotics: a fixed set of tasks on real robots, success rates from a modest number of trials each.