Concept · Chapter 15: Multimodal AI
Vision-Language Models
A vision-language model lets a language model read images by encoding each image into vectors and feeding them into the model alongside text tokens, usually through a small trained connector between a pretrained image encoder and a pretrained LLM.
The problem
LLMs know a great deal about the world in words but receive only text, while image encoders see but cannot talk, follow instructions or reason in language.
The solution
Encode the image with a pretrained vision encoder, map its features into the LLM's input space with a projector (or let the LLM attend to them through cross-attention), then train on image–text data and visual instructions.
The consequence
Chat assistants can answer questions about photos, screenshots, charts and documents, but they inherit the image encoder's blind spots and the LLM's habit of producing fluent text that the image does not support.
You should understand first
- Vectors
- Dot Product
- Embeddings
- Attention
- Probability and Distributions
- Softmax
- Self-Attention
- Multi-Head Attention
- Causal Masking
- Positional Encoding
- Residual Connections
- Layer Normalization
- Feed-Forward Sublayer (MLP)
- The Transformer Block
- Encoder, Decoder & Encoder–Decoder
- Text Embeddings
- Entropy
- Loss Functions
- Cross-Entropy Loss
- Contrastive Learning
- Text as Data
- One-Hot Encoding
- Tokenization
- Tensors and Shapes
- Images as Tensors
- Convolution
- The Turing Test
- Symbolic AI
- Logic and Rules
- Expert Systems
- Knowledge Representation
- The Knowledge-Acquisition Bottleneck
- From Rules to Learning
- Supervised, Unsupervised and Self-Supervised Learning
- Features, Labels and Tasks
- Derivatives and Gradients
- Gradient Descent
- Linear Regression
- Logistic Regression
- The Perceptron
- Activation Functions
- The Artificial Neuron
- Matrix Multiplication
- Multilayer Perceptron (MLP)
- The Chain Rule
- The Forward Pass
- Computational Graphs and Autodiff
- Backpropagation
- Convolutional Neural Networks
- Vision Transformer (ViT)
- Audio and Spectrograms
- Turning Signals into Tokens
- CLIP and Image–Text Embeddings
- Conditional Probability and Bayes' Theorem
- Probability of Sequences
- Language Modeling
- Pretrain, Then Fine-Tune
- Supervised Fine-Tuning
- Vision-Language Models
Three ways to give an LLM eyes
1. Projection into the token stream. Run a pretrained image encoder (often CLIP's) and turn each patch feature into a vector the size of the LLM's token embeddings. Those vectors sit in the sequence like extra words. The first LLaVA used a simple linear layer, a trainable projection matrix, to convert image features into the word-embedding space. Established LLaVA-1.5 replaced it with a two-layer MLP and a CLIP ViT-L encoder at 336 pixels. Established
2. Cross-attention. Leave the LLM's token stream alone and insert new layers through which text tokens attend to image features. Flamingo froze both a vision encoder and a language model, turned image features into a fixed number of visual tokens with a Perceiver Resampler, and conditioned the frozen LM through new gated cross-attention layers whose tanh gates start at zero. Established Starting at zero means the model initially behaves exactly like the original language model.
3. Early fusion. Turn images into discrete tokens in the same vocabulary as text and train one model on mixed sequences from the start. Chameleon encodes a 512 × 512 image into 1,024 tokens from an 8,192-entry codebook, inside a 65,536-token vocabulary shared with text, and can both read and generate images. Established
The first is the most common recipe in open models; the third is the only one of the three that also writes images natively.
The training recipe
A typical projector-style model trains in two stages, much like post-training for text:
- Alignment. Freeze the encoder and the LLM; train only the connector on image–caption pairs so image vectors land where the LLM can use them.
- Visual instruction tuning. Train on conversations about images (questions, answers, descriptions, reasoning), updating the connector and usually the LLM.
LLaVA filtered 595K image–text pairs for stage 1 and used 158K instruction-following samples for stage 2, generated by text-only GPT-4 from each image's captions and bounding boxes. Established The teacher never saw the pixels. That is cheap, and it means some answers describe what the captions said rather than what the image shows, one route to the hallucinations discussed at the end of the chapter.
Tiny example: count the visual tokens
At 336 pixels with 14-pixel patches, one image is 24 × 24 = 576 visual tokens (arithmetic from LLaVA-1.5's encoder). A conversation with four screenshots carries about 2,300 image tokens before any text. Every one of them goes through prefill and sits in the KV cache (Chapter 11). This is why systems downsample, crop into tiles, or compress visual tokens with a resampler: Flamingo's resampler and BLIP-2's Q-Former both hand the LLM a fixed, small number of visual tokens whatever the image size.
What the model sees and what it says
The LLM can only reason about what reached it. If the encoder's resolution blurred a small label, the model may still answer fluently, from the question's wording and its priors. Treat a VLM's description as a claim about the image, not a reading of it. Interpretation Checks that work: ask about details that would be visible only at the input resolution, ask yes/no about objects that are absent, and compare answers on crops.
Where it shows up
Document and chart understanding is the same wiring applied to text-heavy images. Computer-use agents read screenshots with a VLM before acting. Vision-language-action models fine-tune a VLM to output robot actions.
Mini experiment
Take an image-capable assistant and a photo with small printed text. Ask it to transcribe the text, then crop the photo to just the text and ask again. If the answers differ, which one reflects the limit of the encoder's resolution, and which reflects the model filling gaps?
Why should I care?
As a researcher
Most current multimodal research builds on this wiring; choices of encoder, connector, resolution and data decide what the model can perceive and where it hallucinates.
As an engineer
Screenshots, scanned forms, charts and photos become inputs to the same API as text, priced in image tokens, with the same need to check outputs against the source.
Modern systems that depend on it
- document and chart question answering
- computer-use agents that read screens
- vision-language-action models
- multimodal assistants
Historical context
Before
Image captioning and visual question answering used task-specific models trained on labelled datasets; the language model and the vision model were separate systems.
After
One assistant reads images and text together, follows instructions about them, and answers in language, often built from frozen or lightly tuned pretrained parts.
Used today
Multimodal chat assistants, open models in the LLaVA family, document and screenshot understanding, and the perception step of computer-use agents.
What to remember
- Three wirings: project image features into the token sequence (LLaVA), cross-attend to them from the LLM (Flamingo), or tokenize images into the same vocabulary from the start (early fusion, Chameleon).
- LLaVA's first connector was a single trainable matrix; LLaVA-1.5 used a two-layer MLP.
- Training is usually two stages: align the connector on image–caption pairs, then instruction-tune on visual conversations.
- LLaVA's 158K instruction conversations were written by text-only GPT-4 from captions and boxes, not from the images.
- 336-pixel images with 14-pixel patches give 24 × 24 = 576 visual tokens per image.
- The model can only report what its encoder kept: resolution, crop and patch size bound what it can see.
Key papers
Flamingo: a Visual Language Model for Few-Shot Learning
Jean-Baptiste Alayrac, Jeff Donahue et al. · 2022
Flamingo bridged a frozen vision encoder and a frozen language model, and handled interleaved images and text, so a visual language model could learn new tasks from a few examples in the prompt.
How to read it: Figure 4 shows the gated cross-attention block. Its tanh gate starts at zero, so at initialization the model behaves exactly like the frozen language model.
BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
Junnan Li, Dongxu Li et al. · 2023
A cheap bridge between a frozen image encoder and a frozen LLM, showing how much a small trained connector can do.
How to read it: Compare its connector (learned queries that pull a fixed number of tokens out of the image) with LLaVA's (project every patch).
Visual Instruction Tuning
Haotian Liu, Chunyuan Li et al. · 2023
LLaVA showed a simple, open recipe for a visual assistant: a CLIP image encoder, a trained projection into an LLM's word-embedding space, and machine-generated visual instruction data.
How to read it: The two-stage training procedure in Section 4 is the recipe. Section 3 shows that the 'teacher' never saw the images, only their captions and bounding boxes.
Improved Baselines with Visual Instruction Tuning
Haotian Liu, Chunyuan Li et al. · 2023
LLaVA-1.5 showed how far a plain connector goes: a two-layer MLP from CLIP features into the LLM, with the right data, set strong results cheaply.
How to read it: A short note. With 14-pixel patches at 336 pixels, each image becomes 24 × 24 = 576 visual tokens: count them against the context length.
Chameleon: Mixed-Modal Early-Fusion Foundation Models
Chameleon Team · 2024
A clear example of early fusion: images become discrete tokens in the same sequence and vocabulary as text, so one model reads and writes both.
How to read it: The stability section is the practical lesson: mixing modalities in one sequence caused training divergences that needed architectural fixes.