Concept · Chapter 16: Evaluation, Reliability, Safety & Interpretability
Red Teaming
Red teaming means deliberately trying to make a model fail, by people or by other models, to find harmful behaviours before users do and to turn them into tests and training data.
The problem
Ordinary benchmarks sample typical inputs, but the harms that matter most come from rare, creative or adversarial inputs that nobody thought to include.
The solution
Have people with varied expertise attack the model under clear instructions, use language models to generate attacks at scale, classify the failures, and feed them back into training, classifiers and regression tests.
The consequence
Red teaming finds failures and estimates how hard they are to trigger; it cannot prove their absence, so results are reported with the attackers, budget and methods used.
You should understand first
- Text as Data
- Probability and Distributions
- Conditional Probability and Bayes' Theorem
- Probability of Sequences
- Language Modeling
- Vectors
- Dot Product
- Embeddings
- Attention
- Softmax
- Self-Attention
- Multi-Head Attention
- Causal Masking
- Positional Encoding
- Residual Connections
- Layer Normalization
- Feed-Forward Sublayer (MLP)
- The Transformer Block
- Encoder, Decoder & Encoder–Decoder
- Pretrain, Then Fine-Tune
- Entropy
- Loss Functions
- Cross-Entropy Loss
- Supervised Fine-Tuning
- The Turing Test
- Symbolic AI
- Logic and Rules
- Expert Systems
- Knowledge Representation
- The Knowledge-Acquisition Bottleneck
- From Rules to Learning
- Supervised, Unsupervised and Self-Supervised Learning
- Autoregressive Next-Token Prediction
- Decoding: Greedy, Temperature, Top-k, Top-p
- Learning from Comparisons
- Safety Tuning and Refusal
- One-Hot Encoding
- Tokenization
- Chat Templates
- Pretraining at Scale
- GPT-1 → GPT-2 → GPT-3
- In-Context Learning
- System Prompts and Instructions
- Structured Outputs and Constrained Decoding
- Tool Calling
- Guardrails
- Features, Labels and Tasks
- Derivatives and Gradients
- Gradient Descent
- Linear Regression
- Expected Value and Variance
- Sampling and Uncertainty
- Generalization, Overfitting and Underfitting
- Distribution Shift
- Robustness and Prompt Sensitivity
- Jailbreaks and Adversarial Prompts
- Red Teaming
Testing to fail
A benchmark asks how a model behaves on typical inputs. A red team asks how it can be made to behave badly. The term comes from security exercises, where a team plays the attacker.
People
Ganguli and colleagues red-teamed models of 2.7B, 13B and 52B parameters of four kinds and found that models trained with RLHF became harder to red team as they grew, while the other kinds showed no trend with size; they released 38,961 attacks and described their instructions and uncertainties. Established Human red teams find what automated ones miss: context-specific harms, plausible misuse in a domain, and attacks that build over a conversation. Domain experts matter for specialised risks such as medicine or cybersecurity.
Models attacking models
Perez and colleagues used one language model to generate test questions for another and a classifier to detect offensive replies, uncovering tens of thousands of offensive replies in a 280B-parameter chatbot, as well as generated phone numbers and leaked training data. Established Automated red teaming trades depth for scale: thousands of varied attempts, scored by a classifier or judge whose own errors need checking. Optimization-based attacks such as GCG are the extreme case.
Closing the loop
Findings are only useful if they change something:
- Classify each failure: how harmful, how easy to trigger, how general.
- Fix at the right layer: training data, a classifier, a permission, a product change.
- Keep the attacks as a regression suite and rerun them on every new model or prompt.
Mini experiment
For an assistant that answers questions over your company's documents, list five things a red team should try in its first hour. For each, decide whether it is a jailbreak, a prompt injection, a privacy leak or a reliability failure, and which layer should catch it.
What to remember
- Goal: find failures, not score average behaviour.
- Human red teams: diverse expertise, clear instructions, careful handling of harmful data.
- Automated: one LM writes attacks, a classifier or judge scores replies (Perez et al. 2022 found tens of thousands of offensive replies in a 280B chatbot).
- Ganguli et al. 2022: 38,961 red-team attacks released; RLHF models got harder to red team as they scaled.
- Findings become regression tests and training data.
- Absence of findings is weak evidence; report effort and methods.
Key papers
Jailbroken: How Does LLM Safety Training Fail?
Alexander Wei, Nika Haghtalab, Jacob Steinhardt · 2023
Explains why jailbreaks work, in two ideas that still organise the field.
How to read it: Read Section 3 for the two failure modes; the specific attacks have mostly been patched, the reasons have not gone away.
Universal and Transferable Adversarial Attacks on Aligned Language Models
Andy Zou, Zifan Wang et al. · 2023
Showed that jailbreaks can be found automatically by optimization, like adversarial examples for images, and that they transfer between models.
Red Teaming Language Models with Language Models
Ethan Perez, Saffron Huang et al. · 2022
Automated the search for failures: one model writes test cases, another model's replies are classified, at a scale people cannot match.
Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
Deep Ganguli, Liane Lovitt et al. · 2022
A detailed account of human red teaming, with a public dataset of attacks.