Concept · Chapter 15: Multimodal AI
CLIP and Image–Text Embeddings
CLIP trains an image encoder and a text encoder together so that a picture and its caption land on nearby vectors, which lets any written phrase act as a label and gives later models a shared space for pictures and words.
The problem
Vision models learned a fixed list of classes from hand-labelled images, so every new category needed new labels, and their features had no connection to language.
The solution
Collect image–caption pairs from the web and train two encoders contrastively: in each batch, each image must pick out its own caption from all the others, and each caption its own image.
The consequence
Classification becomes 'which caption is closest?', with labels chosen at test time, and generators and vision-language models get an image encoder already aligned with words.
You should understand first
- Vectors
- Dot Product
- Embeddings
- Attention
- Probability and Distributions
- Softmax
- Self-Attention
- Multi-Head Attention
- Causal Masking
- Positional Encoding
- Residual Connections
- Layer Normalization
- Feed-Forward Sublayer (MLP)
- The Transformer Block
- Encoder, Decoder & Encoder–Decoder
- Text Embeddings
- Entropy
- Loss Functions
- Cross-Entropy Loss
- Contrastive Learning
- Text as Data
- One-Hot Encoding
- Tokenization
- Tensors and Shapes
- Images as Tensors
- Convolution
- The Turing Test
- Symbolic AI
- Logic and Rules
- Expert Systems
- Knowledge Representation
- The Knowledge-Acquisition Bottleneck
- From Rules to Learning
- Supervised, Unsupervised and Self-Supervised Learning
- Features, Labels and Tasks
- Derivatives and Gradients
- Gradient Descent
- Linear Regression
- Logistic Regression
- The Perceptron
- Activation Functions
- The Artificial Neuron
- Matrix Multiplication
- Multilayer Perceptron (MLP)
- The Chain Rule
- The Forward Pass
- Computational Graphs and Autodiff
- Backpropagation
- Convolutional Neural Networks
- Vision Transformer (ViT)
- Audio and Spectrograms
- Turning Signals into Tokens
- CLIP and Image–Text Embeddings
Pick the partner out of a line-up
Contrastive learning in Chapter 12 trained a text encoder on question–passage pairs. CLIP makes the two sides different modalities. Take a batch of pictures with their captions. Encode every picture with an image encoder and every caption with a text encoder, normalize each vector, and fill a grid with cosine similarities . Training pushes the diagonal (true pairs) up and everything else down:
The first term asks each picture to find its caption; the second asks each caption to find its picture. The temperature sharpens the softmax. CLIP trained on 400 million image–text pairs collected from the internet, with a minibatch of 32,768 and a learnable temperature initialized to the equivalent of 0.07. Established A bigger batch means more wrong answers to beat in every row, which is part of why contrastive training favours huge batches.
Labels written at test time
Once trained, classification needs no classifier. To sort a photo into dog, cat or bird, embed the texts "a photo of a dog", "a photo of a cat" and "a photo of a bird", and pick the one with the highest similarity to the photo's embedding. The label set is just text, chosen after training.
The CLIP paper reports matching the original ResNet-50's ImageNet accuracy zero-shot, without using any of the 1.28 million labelled training examples that model was trained on. Established The same paper benchmarks over 30 existing datasets, and the model transfers non-trivially to most of them but not uniformly. EstablishedTiny example (lab numbers)
The lab trains a toy CLIP in your browser: 12 × 12 pictures of one coloured shape, captions like "a red circle", a small image network and a bag-of-words text encoder. Four colour–shape combinations are never shown in training, though each colour and each shape appears with other partners.
After 400 batches on the default seed: pictures of seen combinations get the right caption out of 16 about 93% of the time, never-seen combinations about 75%. Asked to choose among the bare words "red", "green", "blue", "yellow", which never appeared alone in training, it is right on every test picture. Across 12 seeds the never-seen score ranges from 55% to 90%: what a model learns about a concept it has seen only in some combinations depends on the run.
What the shared space cannot see
The lab's text encoder adds word vectors, so "a red circle and a blue square" and "a blue circle and a red square" get exactly the same vector. Real CLIP text encoders are Transformers and could see order, yet the problem survives. Yuksekgonul and colleagues tested vision-language models on more than 50,000 cases and found poor relational understanding, mistakes in linking objects to their attributes, and a severe lack of sensitivity to word order. Established Their explanation is the training signal: retrieval on typical datasets succeeds without order information, so the contrastive loss puts little pressure on encoding it. Captions that differ only in order, used as hard negatives, supply that pressure. Interpretation
Other limits follow from the data: concepts absent from the web pairs, small text in images, and the social biases of web captions all carry into the embedding.
Where it shows up
- Vision-language models often use a CLIP-style image encoder as their eyes (vision-language models).
- Image generators condition on text through such embeddings, and DALL·E 2 generated images from CLIP image embeddings (latent diffusion).
- Evaluation: CLIPScore rates how well an image matches text by CLIP similarity (multimodal evaluation).
- Search: the same nearest-neighbour index you built for text in Chapter 12 works for images.
Mini experiment
In the lab, train at τ = 0.1, then retrain at τ = 0.05 and τ = 0.3 and compare the grid's diagonal and the never-seen score. Then choose "Shape words" and find a never-seen combination the model gets wrong. Is the error about colour or about shape, and why might a picture-encoder that saw triangles only in three colours struggle with the fourth?
Why should I care?
As a researcher
CLIP-style encoders are the eyes of many vision-language models and the text conditioning of many image generators, and their limits (word order, counting, small text) propagate into everything built on them.
As an engineer
Image search, deduplication, zero-shot tagging and content filtering can run on CLIP-style embeddings with no labelled training data, using the same vector index as text RAG.
Modern systems that depend on it
- zero-shot image classification
- image–text retrieval
- LLaVA-style vision-language models
- DALL·E 2 and text-conditioned diffusion
- CLIPScore
Historical context
Before
An ImageNet model knew 1,000 fixed classes. A new label meant collecting and labelling new images and retraining.
After
A model trained once on web pairs classifies with any labels you write, retrieves images from text and text from images, and supplies image features already aligned with language.
Used today
Image search and tagging, the vision encoder inside many open vision-language models, text conditioning and evaluation (CLIPScore) for image generators.
What to remember
- Two encoders, one space: matching image–caption pairs get high cosine similarity, mismatched pairs low.
- Loss: a B × B similarity matrix divided by a temperature, cross-entropy along rows and columns (InfoNCE), with the other B − 1 items as free negatives.
- CLIP: 400 million web pairs, batches of 32,768, temperature learned (initialized at 0.07).
- Zero-shot: embed 'a photo of a ⟨label⟩' for each label and pick the closest; it matched the original ResNet-50's ImageNet accuracy without its 1.28M labelled images.
- Embeddings of a whole caption can lose word order: 'red circle, blue square' vs 'blue circle, red square'.
- Zero-shot only reaches concepts the training pairs covered; it is not knowledge of everything.
Key papers
Learning Transferable Visual Models From Natural Language Supervision
Alec Radford, Jong Wook Kim et al. · 2021 · ICML 2021
CLIP learned a shared space for images and text from web captions — the backbone of much multimodal AI and text-to-image generation.
When and why vision-language models behave like bags-of-words, and what to do about it?
Mert Yuksekgonul, Federico Bianchi et al. · 2022
Measured, at scale, that contrastive image–text models often ignore word order and mis-bind attributes, and explained why their training and benchmarks let them.
How to read it: The key argument is about the training objective: if shuffled captions still retrieve the right image, the contrastive loss never needed word order. Hard negatives that differ only in order fix the incentive.
CLIPScore: A Reference-free Evaluation Metric for Image Captioning
Jack Hessel, Ari Holtzman et al. · 2021
Using CLIP's image–text similarity as a metric became standard for captioning and text-to-image evaluation, along with its blind spots.
How to read it: Read the case studies at the end: a metric built on one model inherits that model's blind spots.
Hierarchical Text-Conditional Image Generation with CLIP Latents
Aditya Ramesh, Prafulla Dhariwal et al. · 2022
DALL·E 2 generated images through CLIP's embedding space, a direct demonstration that the shared space CLIP learned can be run backwards.
How to read it: Look at the variations figures: the embedding fixes the meaning and style, and everything the embedding leaves out changes.
Watch
3Blue1Brown
But how do AI images and videos actually work? | Guest video by Welch Labs
A guest video by Welch Labs: animated, careful and math-first, the clearest single explanation of how CLIP and diffusion fit together in text-to-image models.