Skip to content
Road to Intelligence

Concept · Chapter 16: Evaluation, Reliability, Safety & Interpretability

Memorization and Training-Data Extraction

Should knowUnderstand11 minDifficulty

Language models store some of their training data verbatim, and an attacker who can only query the model can sometimes get it back out, including personal information that appeared in a single document.

The problem

Training data scraped from the web or collected from users contains personal and confidential text, and it was assumed that a model trained on billions of documents learns patterns rather than storing individual ones.

The solution

Measure memorization with extraction and membership-inference attacks, reduce it by deduplicating and filtering training data, and where guarantees are needed train with differential privacy, which bounds how much any single example can influence the model.

The consequence

Memorization is a measurable property that grows with model size and repetition; what goes into training data can come out, so data choices are privacy choices.

What goes in can come out

Next-token training rewards predicting the training text exactly. For text seen many times, such as a famous poem or a licence, exact recall is the best prediction. The surprise was how much else is stored.

Carlini and colleagues generated samples from GPT-2, ranked them by signals of memorization, and confirmed 604 unique memorized training examples among 1,800 candidates, including names, phone numbers, email addresses, code and 128-bit UUIDs, some of which appeared in just one document; larger models were more vulnerable. Established

What makes memorization more likely

A follow-up study measured three log-linear relationships: memorization grows with model capacity, with the number of times an example is duplicated in training, and with the number of tokens of context used to prompt for it (Carlini et al. 2022). Established That ties directly to deduplication in Chapter 9: repeated text is both wasted compute and a privacy risk. The same study found that deduplication helps but does not prevent leakage, because even a few duplicates raise memorization. Established

Aligned models still remember

Nasr and colleagues asked ChatGPT to repeat a single word such as "poem" forever; after many repetitions it diverged and emitted training data at 150 times its normal rate, and about 200 US dollars of queries yielded over 10,000 unique memorized examples. Established Post-training changed what the model usually says, not what it had stored. Interpretation

Two attacker questions

  • Extraction: can I make the model output its training data?
  • Membership inference: given a record, can I tell whether it was in the training set? Shokri and colleagues did this by training "shadow" models to learn how a model behaves differently on its own training inputs. Established In a medical dataset, membership alone is sensitive.

The same questions apply to images: diffusion models have reproduced training images (Chapter 15).

Defences

  1. Data: deduplicate, filter personal data and secrets before training; do not train on what you cannot afford to leak.
  2. Differential privacy: DP-SGD clips each example's gradient, adds Gaussian noise to the sum and tracks a formal privacy budget across training. Established It bounds how much one example can change the model, at a cost in accuracy that grows with the strength of the guarantee.
  3. Output checks: detect and block known secrets or long verbatim runs.

Tiny example

Suppose a support chatbot is fine-tuned on 50,000 past tickets, and one customer's address appears in 40 of them (they wrote in often). Each duplicate raises the chance that a prompt such as "Ship to: [their name]," is completed with the real address. Deduplicating tickets and replacing personal fields with placeholders before fine-tuning removes most of that risk at no cost to what the bot needs to learn.

Mini experiment

List the data sources your team would fine-tune on. For each: does it contain personal or confidential text, how often is it repeated, and who could query the resulting model? Which sources would you scrub, deduplicate or leave out?

What to remember

  • Extraction: generate many samples, rank by signs of memorization, verify against the training data.
  • Carlini et al. 2020: 604 verbatim training examples recovered from GPT-2, including names, phone numbers and emails; larger models were more vulnerable.
  • Memorization grows log-linearly with model size, number of duplicates and prompt length (Carlini et al. 2022).
  • Alignment hides, not removes: a 'repeat this word forever' attack made ChatGPT emit training data 150× more often.
  • Membership inference: was this record in the training set?
  • Defences: deduplicate, filter personal data, DP-SGD (clip per-example gradients, add noise) where guarantees matter.

Key papers

Important

Extracting Training Data from Diffusion Models

Nicholas Carlini, Jamie Hayes et al. · 2023

Showed that image generators can reproduce individual training images, which matters for privacy and copyright and for what 'generating new images' means.

How to read it: Look at how 'memorized' is defined before reading the counts, and note the role of duplicated training images.

~40 min readarXiv:2301.13188✓ verified 2026-10-07
Essential

Extracting Training Data from Large Language Models

Nicholas Carlini, Florian Tramer et al. · 2020

The first clear demonstration that a public language model can be queried to reproduce individual training documents, including personal data.

~35 min readarXiv:2012.07805✓ verified 2026-10-07
Optional

Deep Learning with Differential Privacy

Martín Abadi, Andy Chu et al. · 2016

The standard way to train a network with a mathematical limit on how much any one training example can influence it.

~30 min readarXiv:1607.00133✓ verified 2026-10-07