Skip to content
Road to Intelligence

Concept · Chapter 16: Evaluation, Reliability, Safety & Interpretability

Red Teaming

Should knowUnderstand9 minDifficulty

Red teaming means deliberately trying to make a model fail, by people or by other models, to find harmful behaviours before users do and to turn them into tests and training data.

The problem

Ordinary benchmarks sample typical inputs, but the harms that matter most come from rare, creative or adversarial inputs that nobody thought to include.

The solution

Have people with varied expertise attack the model under clear instructions, use language models to generate attacks at scale, classify the failures, and feed them back into training, classifiers and regression tests.

The consequence

Red teaming finds failures and estimates how hard they are to trigger; it cannot prove their absence, so results are reported with the attackers, budget and methods used.

Testing to fail

A benchmark asks how a model behaves on typical inputs. A red team asks how it can be made to behave badly. The term comes from security exercises, where a team plays the attacker.

People

Ganguli and colleagues red-teamed models of 2.7B, 13B and 52B parameters of four kinds and found that models trained with RLHF became harder to red team as they grew, while the other kinds showed no trend with size; they released 38,961 attacks and described their instructions and uncertainties. Established Human red teams find what automated ones miss: context-specific harms, plausible misuse in a domain, and attacks that build over a conversation. Domain experts matter for specialised risks such as medicine or cybersecurity.

Models attacking models

Perez and colleagues used one language model to generate test questions for another and a classifier to detect offensive replies, uncovering tens of thousands of offensive replies in a 280B-parameter chatbot, as well as generated phone numbers and leaked training data. Established Automated red teaming trades depth for scale: thousands of varied attempts, scored by a classifier or judge whose own errors need checking. Optimization-based attacks such as GCG are the extreme case.

Closing the loop

Findings are only useful if they change something:

  1. Classify each failure: how harmful, how easy to trigger, how general.
  2. Fix at the right layer: training data, a classifier, a permission, a product change.
  3. Keep the attacks as a regression suite and rerun them on every new model or prompt.

Mini experiment

For an assistant that answers questions over your company's documents, list five things a red team should try in its first hour. For each, decide whether it is a jailbreak, a prompt injection, a privacy leak or a reliability failure, and which layer should catch it.

What to remember

  • Goal: find failures, not score average behaviour.
  • Human red teams: diverse expertise, clear instructions, careful handling of harmful data.
  • Automated: one LM writes attacks, a classifier or judge scores replies (Perez et al. 2022 found tens of thousands of offensive replies in a 280B chatbot).
  • Ganguli et al. 2022: 38,961 red-team attacks released; RLHF models got harder to red team as they scaled.
  • Findings become regression tests and training data.
  • Absence of findings is weak evidence; report effort and methods.

Key papers

Essential

Jailbroken: How Does LLM Safety Training Fail?

Alexander Wei, Nika Haghtalab, Jacob Steinhardt · 2023

Explains why jailbreaks work, in two ideas that still organise the field.

How to read it: Read Section 3 for the two failure modes; the specific attacks have mostly been patched, the reasons have not gone away.

~30 min readarXiv:2307.02483✓ verified 2026-10-07
Important

Red Teaming Language Models with Language Models

Ethan Perez, Saffron Huang et al. · 2022

Automated the search for failures: one model writes test cases, another model's replies are classified, at a scale people cannot match.

~30 min readarXiv:2202.03286✓ verified 2026-10-07