Skip to content
Road to Intelligence

Concept · Chapter 9: How an LLM Is Actually Built

Benchmark Contamination

Should knowUnderstand8 minDifficulty

Contamination happens when benchmark test questions (or their answers) end up in a model's training data, so its score partly measures memory rather than ability.

The problem

Web-scale training sets are scraped from the same internet where benchmarks, their solutions and discussions of them are published.

The solution

Search the training data for overlaps with evaluation sets and remove them (decontamination), report what was found, and prefer fresh or held-out evaluations.

The consequence

Benchmark numbers from web-trained models always carry some doubt; careful reports measure and disclose overlap rather than assuming it away.

Leakage, at the scale of the internet

Chapter 3's rule was simple: never let the test set influence training. With web-scale data it is hard to keep. Benchmarks are published online, copied into tutorials and forums, and solved in blog posts, and the crawler picks all of it up.

GPT-3's authors searched their training data for overlaps with the benchmarks they reported, but a bug in the filtering left some overlaps in, and the cost of training made it infeasible to retrain; they analysed the impact instead Established. An audit of the C4 dataset found examples from benchmark NLP datasets in its text Established, and Lee and colleagues found that over 4% of the validation sets of standard datasets overlapped with their training sets Established.

What teams do about it

  • Decontaminate: look for long n-gram matches between training documents and evaluation items, and drop the documents.
  • Measure and report how much of each benchmark overlapped, and compare scores on clean and contaminated subsets.
  • Hold data back: Llama 3's team deliberately left the training sets of commonly used benchmarks out of its final high-quality "annealing" data, so they could assess few-shot ability honestly Established.
  • Use newer evaluations, written after the training data was collected.

What to remember

  • Contamination is data leakage (Chapter 3) at web scale.
  • GPT-3's authors found overlaps but a bug left some in, and retraining was too costly.
  • An audit of C4 found examples from NLP benchmarks inside it.
  • n-gram overlap checks catch copies, not paraphrases or translations.

Key papers

Essential

Language Models are Few-Shot Learners

Tom B. Brown, Benjamin Mann et al. · 2020 · NeurIPS 2020

GPT-3 (175B parameters) showed that a large enough language model can perform new tasks from a few examples in its prompt, without any gradient updates.

How to read it: 75 pages. Sections 1–2 and Figure 1.2 carry the core idea; Section 6 on broader impacts is worth reading too.

~1 h 30 min readarXiv:2005.14165✓ verified 2026-09-26
Important

Deduplicating Training Data Makes Language Models Better

Katherine Lee, Daphne Ippolito et al. · 2021

Showed that standard training sets are full of duplicates, and that removing them reduces memorisation and leaks between training and test data.

~30 min readarXiv:2107.06499✓ verified 2026-10-04