Concept · Chapter 12: Embeddings, RAG & the LLM Application Stack
Contrastive Learning
Contrastive learning trains an encoder by pulling the vectors of matching pairs (a question and its answer) together and pushing non-matching ones apart, usually by asking the model to pick the true partner out of a batch with a softmax.
The problem
An embedding is only useful if 'close' means 'relevant', and nothing in ordinary pretraining guarantees that.
The solution
Collect pairs that should match; for each query, score its true partner against the other items in the batch and train with cross-entropy so the partner wins.
The consequence
Cheap supervision at scale: any naturally paired text (titles and articles, questions and answers) becomes training data, and the other items in a batch serve as free negative examples. CLIP (Chapter 15) uses the same idea for images and text.
You should understand first
- Vectors
- Dot Product
- Embeddings
- Attention
- Probability and Distributions
- Softmax
- Self-Attention
- Multi-Head Attention
- Causal Masking
- Positional Encoding
- Residual Connections
- Layer Normalization
- Feed-Forward Sublayer (MLP)
- The Transformer Block
- Encoder, Decoder & Encoder–Decoder
- Text Embeddings
- Entropy
- Loss Functions
- Cross-Entropy Loss
- Contrastive Learning
Pick the partner out of a line-up
Give the encoder a batch of question–passage pairs. For question , its passage is the right answer and the other passages are wrong answers, for free. Score every passage against the question with a dot product, turn the scores into probabilities with a softmax, and use cross-entropy to raise the probability of the true partner:
is a temperature. This form of loss, a softmax that picks the positive sample out of a set of negatives, was named InfoNCE by van den Oord and colleagues Established; it is the same "pick the right word out of the vocabulary" cross-entropy a language model uses, except the "vocabulary" is the batch.
Tiny example. A question scores 0.9 with its passage and 0.2, 0.1 and 0.4 with three others. With τ = 0.1, the softmax gives the true passage , so the loss is about 0.01: little to learn. If it had scored 0.4 against a negative's 0.5, the true passage would get under half the probability and the gradient would push the two apart.
In-batch and hard negatives
Dense Passage Retrieval trained its question and passage encoders with in-batch negatives, reusing the other questions' passages in a batch as negatives, and found it an effective and efficient training strategy Established. Bigger batches mean more negatives per step.
Random negatives are easy to reject. Hard negatives, passages that look relevant but aren't (for example a high-scoring BM25 result that doesn't contain the answer), force the model to learn finer distinctions. Modern embedding models combine huge numbers of weakly paired web texts with smaller sets of carefully labelled pairs and hard negatives.
What to remember
- Training data is pairs that should match; negatives are everything else in the batch.
- Loss: softmax over similarities, cross-entropy on the true partner (InfoNCE).
- A temperature sharpens or softens the softmax.
- Hard negatives (similar but wrong) teach finer distinctions than random ones.
Key papers
Representation Learning with Contrastive Predictive Coding
Aaron van den Oord, Yazhe Li, Oriol Vinyals · 2018
Introduced the InfoNCE loss: pick the true partner out of a set of negatives with a softmax, the objective most embedding models are trained with.
Dense Passage Retrieval for Open-Domain Question Answering
Vladimir Karpukhin, Barlas Oğuz et al. · 2020
Showed that learned embeddings alone can beat BM25 for finding answer passages, which made dense retrieval the default first stage of RAG.
How to read it: Section 3 (the dual encoder and in-batch negatives) is the core; Section 4.1 describes splitting Wikipedia into 21 million 100-word passages.
Text Embeddings by Weakly-Supervised Contrastive Pre-training
Liang Wang, Nan Yang et al. · 2022
E5 showed the modern recipe for general-purpose embedding models: contrastive pretraining on huge numbers of naturally occurring text pairs, then fine-tuning.