Concept · Chapter 16: Evaluation, Reliability, Safety & Interpretability
Jailbreaks and Adversarial Prompts
A jailbreak is an input crafted to make a safety-trained model do what its training was meant to prevent, and it works because safety training is a learned tendency competing with other tendencies, not a rule the model must obey.
The problem
Models are trained to decline harmful requests, but users and attackers keep finding wordings, encodings, role-plays and optimized strings that get around the refusal.
The solution
Understand why attacks work (competing objectives, safety training that does not generalize as far as capability), test with red teaming and automated attacks, and defend in layers: training, input and output classifiers, limited permissions, and monitoring.
The consequence
No model should be assumed jailbreak-proof; applications are designed so that a successful jailbreak gains little, and safety claims are stated against specific attacks.
You should understand first
- Text as Data
- Probability and Distributions
- Conditional Probability and Bayes' Theorem
- Probability of Sequences
- Language Modeling
- Vectors
- Dot Product
- Embeddings
- Attention
- Softmax
- Self-Attention
- Multi-Head Attention
- Causal Masking
- Positional Encoding
- Residual Connections
- Layer Normalization
- Feed-Forward Sublayer (MLP)
- The Transformer Block
- Encoder, Decoder & Encoder–Decoder
- Pretrain, Then Fine-Tune
- Entropy
- Loss Functions
- Cross-Entropy Loss
- Supervised Fine-Tuning
- The Turing Test
- Symbolic AI
- Logic and Rules
- Expert Systems
- Knowledge Representation
- The Knowledge-Acquisition Bottleneck
- From Rules to Learning
- Supervised, Unsupervised and Self-Supervised Learning
- Autoregressive Next-Token Prediction
- Decoding: Greedy, Temperature, Top-k, Top-p
- Learning from Comparisons
- Safety Tuning and Refusal
- One-Hot Encoding
- Tokenization
- Chat Templates
- Pretraining at Scale
- GPT-1 → GPT-2 → GPT-3
- In-Context Learning
- System Prompts and Instructions
- Structured Outputs and Constrained Decoding
- Tool Calling
- Guardrails
- Features, Labels and Tasks
- Derivatives and Gradients
- Gradient Descent
- Linear Regression
- Expected Value and Variance
- Sampling and Uncertainty
- Generalization, Overfitting and Underfitting
- Distribution Shift
- Robustness and Prompt Sensitivity
- Jailbreaks and Adversarial Prompts
Training is a tendency, not a rule
Safety tuning teaches a model to decline certain requests. It does so with examples and preferences, the same way the model learned everything else. The result is a strong tendency, not a hard-coded check. A jailbreak is any input that makes another tendency win.
Two reasons attacks work
Wei, Haghtalab and Steinhardt identified two failure modes of safety training. Competing objectives: the model's capabilities and its goals of being helpful and following instructions conflict with its safety goal, for example when told to begin its answer with "Absolutely! Here's". Mismatched generalization: safety training does not reach inputs where the model's capabilities still work, for example a request written in Base64, which the model can decode. Established Attacks designed around these failure modes succeeded on every prompt in a collection of unsafe requests against GPT-4 and Claude v1.3, despite extensive safety training. Established Their conclusion is the important one: safety mechanisms need to be as capable as the model they protect. A model that can read base64 needs safety behaviour that also reads base64. InterpretationAutomated attacks
Zou and colleagues found adversarial suffixes automatically with Greedy Coordinate Gradient search: optimize 20 tokens appended to a request so the model's most likely reply begins affirmatively. It succeeded on 88% of target harmful strings for Vicuna-7B and on 100% of harmful behaviours (88% for Llama-2-7B-Chat), and suffixes optimized on open models transferred to ChatGPT, Bard and Claude. Established The suffixes look like gibberish, which is the adversarial example story of 2013 moved to text.
Jailbreak or prompt injection?
The two are often confused:
| Jailbreak | Prompt injection | |
|---|---|---|
| Who attacks | The user | A third party, via content the app reads |
| Target | The model's own rules | The application's instructions and permissions |
| Example | Role-play to get refused content | A web page that tells the assistant to email your files |
Prompt injection is usually the bigger risk for applications, because the user is the victim, not the attacker. Greshake and colleagues showed instructions planted in retrieved content taking over LLM-integrated applications. Established
Defending
No current defence makes a model reliably resistant to adaptive attackers, so defence is layered. Interpretation- Training: refusals and safe completions, including on encodings and role-plays found by red teaming.
- Classifiers on inputs and outputs, trained separately from the model.
- Least privilege: whatever the model says, it can only do what its tools and permissions allow (guardrails).
- Monitoring: detect and respond to repeated or novel attacks.
Mini experiment
Think of an assistant you would deploy. Write down the worst thing a fully jailbroken model could cause through it: what can it read, write, send or spend? Each item on that list is a permission to remove or gate, whatever the model's refusal rate.
Why should I care?
As a researcher
Jailbreaks are evidence about what safety training actually changes inside a model, and they motivate work on robustness, interpretability-based monitoring and adversarial training.
As an engineer
If your app's safety depends on the model refusing, assume a determined user will get past it; put the real limits in code, permissions and output checks.
Modern systems that depend on it
- red-teaming programs
- input and output safety classifiers
- adversarial training data
- agent permission design
Historical context
Before
Safety tuning was evaluated on straightforward harmful requests, and refusals on those were taken as evidence of safety.
After
Safety is evaluated against adaptive attacks, including automated ones, and treated as one layer in a system that limits what any single model output can do.
Used today
Red-team evaluations before releases, attack benchmarks, safety classifiers on inputs and outputs, and permission boundaries in agents.
What to remember
- Jailbreak: get a safety-trained model to produce what it was trained to refuse. Prompt injection: get an application to follow instructions hidden in data.
- Competing objectives: helpfulness or instruction-following pulls against refusing (role-play, 'begin your answer with Absolutely! Here is').
- Mismatched generalization: capabilities reach inputs that safety training never covered (encodings, other languages).
- Wei et al. 2023: attacks built on these two ideas succeeded on every prompt in a set of unsafe requests against GPT-4 and Claude v1.3.
- GCG (2023): an automatically optimized 20-token suffix; 88% on harmful strings for Vicuna-7B, transferred to other models.
- Defence in depth: training + classifiers + least privilege + monitoring. Limit the damage, not just the attempts.
Key papers
Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection
Kai Greshake, Sahar Abdelnabi et al. · 2023
Showed that text a model retrieves (a web page, an email, a document) can carry instructions an attacker planted, so retrieval and tools are a security boundary.
Jailbroken: How Does LLM Safety Training Fail?
Alexander Wei, Nika Haghtalab, Jacob Steinhardt · 2023
Explains why jailbreaks work, in two ideas that still organise the field.
How to read it: Read Section 3 for the two failure modes; the specific attacks have mostly been patched, the reasons have not gone away.
Universal and Transferable Adversarial Attacks on Aligned Language Models
Andy Zou, Zifan Wang et al. · 2023
Showed that jailbreaks can be found automatically by optimization, like adversarial examples for images, and that they transfer between models.
Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
Deep Ganguli, Liane Lovitt et al. · 2022
A detailed account of human red teaming, with a public dataset of attacks.
Watch
Andrej Karpathy
[1hr Talk] Intro to Large Language Models
A clear one-hour overview of what LLMs are, how they are trained, and where they're going — good orientation for Part III.