Concept · Chapter 16: Evaluation, Reliability, Safety & Interpretability
Memorization and Training-Data Extraction
Language models store some of their training data verbatim, and an attacker who can only query the model can sometimes get it back out, including personal information that appeared in a single document.
The problem
Training data scraped from the web or collected from users contains personal and confidential text, and it was assumed that a model trained on billions of documents learns patterns rather than storing individual ones.
The solution
Measure memorization with extraction and membership-inference attacks, reduce it by deduplicating and filtering training data, and where guarantees are needed train with differential privacy, which bounds how much any single example can influence the model.
The consequence
Memorization is a measurable property that grows with model size and repetition; what goes into training data can come out, so data choices are privacy choices.
You should understand first
- Text as Data
- Probability and Distributions
- Conditional Probability and Bayes' Theorem
- Probability of Sequences
- Language Modeling
- Vectors
- Dot Product
- Embeddings
- Attention
- Softmax
- Self-Attention
- Causal Masking
- Entropy
- Loss Functions
- Cross-Entropy Loss
- Autoregressive Next-Token Prediction
- Pretraining at Scale
- One-Hot Encoding
- Tokenization
- Building a Pretraining Dataset
- Deduplication and MinHash
- The Turing Test
- Symbolic AI
- Logic and Rules
- Expert Systems
- Knowledge Representation
- The Knowledge-Acquisition Bottleneck
- From Rules to Learning
- Supervised, Unsupervised and Self-Supervised Learning
- Features, Labels and Tasks
- Derivatives and Gradients
- Gradient Descent
- Linear Regression
- Expected Value and Variance
- Sampling and Uncertainty
- Generalization, Overfitting and Underfitting
- Memorization and Training-Data Extraction
What goes in can come out
Next-token training rewards predicting the training text exactly. For text seen many times, such as a famous poem or a licence, exact recall is the best prediction. The surprise was how much else is stored.
Carlini and colleagues generated samples from GPT-2, ranked them by signals of memorization, and confirmed 604 unique memorized training examples among 1,800 candidates, including names, phone numbers, email addresses, code and 128-bit UUIDs, some of which appeared in just one document; larger models were more vulnerable. EstablishedWhat makes memorization more likely
A follow-up study measured three log-linear relationships: memorization grows with model capacity, with the number of times an example is duplicated in training, and with the number of tokens of context used to prompt for it (Carlini et al. 2022). Established That ties directly to deduplication in Chapter 9: repeated text is both wasted compute and a privacy risk. The same study found that deduplication helps but does not prevent leakage, because even a few duplicates raise memorization. Established
Aligned models still remember
Nasr and colleagues asked ChatGPT to repeat a single word such as "poem" forever; after many repetitions it diverged and emitted training data at 150 times its normal rate, and about 200 US dollars of queries yielded over 10,000 unique memorized examples. Established Post-training changed what the model usually says, not what it had stored. InterpretationTwo attacker questions
- Extraction: can I make the model output its training data?
- Membership inference: given a record, can I tell whether it was in the training set? Shokri and colleagues did this by training "shadow" models to learn how a model behaves differently on its own training inputs. Established In a medical dataset, membership alone is sensitive.
The same questions apply to images: diffusion models have reproduced training images (Chapter 15).
Defences
- Data: deduplicate, filter personal data and secrets before training; do not train on what you cannot afford to leak.
- Differential privacy: DP-SGD clips each example's gradient, adds Gaussian noise to the sum and tracks a formal privacy budget across training. Established It bounds how much one example can change the model, at a cost in accuracy that grows with the strength of the guarantee.
- Output checks: detect and block known secrets or long verbatim runs.
Tiny example
Suppose a support chatbot is fine-tuned on 50,000 past tickets, and one customer's address appears in 40 of them (they wrote in often). Each duplicate raises the chance that a prompt such as "Ship to: [their name]," is completed with the real address. Deduplicating tickets and replacing personal fields with placeholders before fine-tuning removes most of that risk at no cost to what the bot needs to learn.
Mini experiment
List the data sources your team would fine-tune on. For each: does it contain personal or confidential text, how often is it repeated, and who could query the resulting model? Which sources would you scrub, deduplicate or leave out?
What to remember
- Extraction: generate many samples, rank by signs of memorization, verify against the training data.
- Carlini et al. 2020: 604 verbatim training examples recovered from GPT-2, including names, phone numbers and emails; larger models were more vulnerable.
- Memorization grows log-linearly with model size, number of duplicates and prompt length (Carlini et al. 2022).
- Alignment hides, not removes: a 'repeat this word forever' attack made ChatGPT emit training data 150× more often.
- Membership inference: was this record in the training set?
- Defences: deduplicate, filter personal data, DP-SGD (clip per-example gradients, add noise) where guarantees matter.
Key papers
Extracting Training Data from Diffusion Models
Nicholas Carlini, Jamie Hayes et al. · 2023
Showed that image generators can reproduce individual training images, which matters for privacy and copyright and for what 'generating new images' means.
How to read it: Look at how 'memorized' is defined before reading the counts, and note the role of duplicated training images.
Extracting Training Data from Large Language Models
Nicholas Carlini, Florian Tramer et al. · 2020
The first clear demonstration that a public language model can be queried to reproduce individual training documents, including personal data.
Quantifying Memorization Across Neural Language Models
Nicholas Carlini, Daphne Ippolito et al. · 2022
Turned memorization into three measured trends, all pointing the wrong way as models grow.
Scalable Extraction of Training Data from (Production) Language Models
Milad Nasr, Nicholas Carlini et al. · 2023
Showed that alignment hides memorization rather than removing it, with a strikingly simple attack on ChatGPT.
Deep Learning with Differential Privacy
Martín Abadi, Andy Chu et al. · 2016
The standard way to train a network with a mathematical limit on how much any one training example can influence it.
Membership Inference Attacks against Machine Learning Models
Reza Shokri, Marco Stronati et al. · 2016
Defined the basic privacy question for any trained model: was this record in the training data?