Skip to content
Road to Intelligence

Concept · Chapter 12: Embeddings, RAG & the LLM Application Stack

Contrastive Learning

Should knowUnderstand10 minDifficulty

Contrastive learning trains an encoder by pulling the vectors of matching pairs (a question and its answer) together and pushing non-matching ones apart, usually by asking the model to pick the true partner out of a batch with a softmax.

The problem

An embedding is only useful if 'close' means 'relevant', and nothing in ordinary pretraining guarantees that.

The solution

Collect pairs that should match; for each query, score its true partner against the other items in the batch and train with cross-entropy so the partner wins.

The consequence

Cheap supervision at scale: any naturally paired text (titles and articles, questions and answers) becomes training data, and the other items in a batch serve as free negative examples. CLIP (Chapter 15) uses the same idea for images and text.

Pick the partner out of a line-up

Give the encoder a batch of BB question–passage pairs. For question ii, its passage is the right answer and the other B−1B - 1 passages are wrong answers, for free. Score every passage against the question with a dot product, turn the scores into probabilities with a softmax, and use cross-entropy to raise the probability of the true partner:

Li=−log⁡exp⁡(qi⋅pi/τ)∑j=1Bexp⁡(qi⋅pj/τ).\mathcal{L}_i = -\log \frac{\exp(\mathbf{q}_i \cdot \mathbf{p}_i / \tau)}{\sum_{j=1}^{B} \exp(\mathbf{q}_i \cdot \mathbf{p}_j / \tau)}.

τ\tau is a temperature. This form of loss, a softmax that picks the positive sample out of a set of negatives, was named InfoNCE by van den Oord and colleagues Established; it is the same "pick the right word out of the vocabulary" cross-entropy a language model uses, except the "vocabulary" is the batch.

Tiny example. A question scores 0.9 with its passage and 0.2, 0.1 and 0.4 with three others. With τ = 0.1, the softmax gives the true passage e9/(e9+e2+e1+e4)≈0.99e^{9} / (e^{9} + e^{2} + e^{1} + e^{4}) \approx 0.99, so the loss is about 0.01: little to learn. If it had scored 0.4 against a negative's 0.5, the true passage would get under half the probability and the gradient would push the two apart.

In-batch and hard negatives

Dense Passage Retrieval trained its question and passage encoders with in-batch negatives, reusing the other questions' passages in a batch as negatives, and found it an effective and efficient training strategy Established. Bigger batches mean more negatives per step.

Random negatives are easy to reject. Hard negatives, passages that look relevant but aren't (for example a high-scoring BM25 result that doesn't contain the answer), force the model to learn finer distinctions. Modern embedding models combine huge numbers of weakly paired web texts with smaller sets of carefully labelled pairs and hard negatives.

What to remember

  • Training data is pairs that should match; negatives are everything else in the batch.
  • Loss: softmax over similarities, cross-entropy on the true partner (InfoNCE).
  • A temperature sharpens or softens the softmax.
  • Hard negatives (similar but wrong) teach finer distinctions than random ones.

Key papers

Optional

Representation Learning with Contrastive Predictive Coding

Aaron van den Oord, Yazhe Li, Oriol Vinyals · 2018

Introduced the InfoNCE loss: pick the true partner out of a set of negatives with a softmax, the objective most embedding models are trained with.

~40 min readarXiv:1807.03748✓ verified 2026-10-05
Essential

Dense Passage Retrieval for Open-Domain Question Answering

Vladimir Karpukhin, Barlas Oğuz et al. · 2020

Showed that learned embeddings alone can beat BM25 for finding answer passages, which made dense retrieval the default first stage of RAG.

How to read it: Section 3 (the dual encoder and in-batch negatives) is the core; Section 4.1 describes splitting Wikipedia into 21 million 100-word passages.

~40 min readarXiv:2004.04906✓ verified 2026-10-05
Important

Text Embeddings by Weakly-Supervised Contrastive Pre-training

Liang Wang, Nan Yang et al. · 2022

E5 showed the modern recipe for general-purpose embedding models: contrastive pretraining on huge numbers of naturally occurring text pairs, then fine-tuning.

~35 min readarXiv:2212.03533✓ verified 2026-10-05