Skip to content
Road to Intelligence

Concept · Chapter 16: Evaluation, Reliability, Safety & Interpretability

Jailbreaks and Adversarial Prompts

Must knowKnow well13 minDifficulty

A jailbreak is an input crafted to make a safety-trained model do what its training was meant to prevent, and it works because safety training is a learned tendency competing with other tendencies, not a rule the model must obey.

The problem

Models are trained to decline harmful requests, but users and attackers keep finding wordings, encodings, role-plays and optimized strings that get around the refusal.

The solution

Understand why attacks work (competing objectives, safety training that does not generalize as far as capability), test with red teaming and automated attacks, and defend in layers: training, input and output classifiers, limited permissions, and monitoring.

The consequence

No model should be assumed jailbreak-proof; applications are designed so that a successful jailbreak gains little, and safety claims are stated against specific attacks.

Training is a tendency, not a rule

Safety tuning teaches a model to decline certain requests. It does so with examples and preferences, the same way the model learned everything else. The result is a strong tendency, not a hard-coded check. A jailbreak is any input that makes another tendency win.

Two reasons attacks work

Wei, Haghtalab and Steinhardt identified two failure modes of safety training. Competing objectives: the model's capabilities and its goals of being helpful and following instructions conflict with its safety goal, for example when told to begin its answer with "Absolutely! Here's". Mismatched generalization: safety training does not reach inputs where the model's capabilities still work, for example a request written in Base64, which the model can decode. Established Attacks designed around these failure modes succeeded on every prompt in a collection of unsafe requests against GPT-4 and Claude v1.3, despite extensive safety training. Established Their conclusion is the important one: safety mechanisms need to be as capable as the model they protect. A model that can read base64 needs safety behaviour that also reads base64. Interpretation

Automated attacks

Zou and colleagues found adversarial suffixes automatically with Greedy Coordinate Gradient search: optimize 20 tokens appended to a request so the model's most likely reply begins affirmatively. It succeeded on 88% of target harmful strings for Vicuna-7B and on 100% of harmful behaviours (88% for Llama-2-7B-Chat), and suffixes optimized on open models transferred to ChatGPT, Bard and Claude. Established The suffixes look like gibberish, which is the adversarial example story of 2013 moved to text.

Jailbreak or prompt injection?

The two are often confused:

JailbreakPrompt injection
Who attacksThe userA third party, via content the app reads
TargetThe model's own rulesThe application's instructions and permissions
ExampleRole-play to get refused contentA web page that tells the assistant to email your files

Prompt injection is usually the bigger risk for applications, because the user is the victim, not the attacker. Greshake and colleagues showed instructions planted in retrieved content taking over LLM-integrated applications. Established

Defending

No current defence makes a model reliably resistant to adaptive attackers, so defence is layered. Interpretation
  1. Training: refusals and safe completions, including on encodings and role-plays found by red teaming.
  2. Classifiers on inputs and outputs, trained separately from the model.
  3. Least privilege: whatever the model says, it can only do what its tools and permissions allow (guardrails).
  4. Monitoring: detect and respond to repeated or novel attacks.

Mini experiment

Think of an assistant you would deploy. Write down the worst thing a fully jailbroken model could cause through it: what can it read, write, send or spend? Each item on that list is a permission to remove or gate, whatever the model's refusal rate.

Why should I care?

As a researcher

Jailbreaks are evidence about what safety training actually changes inside a model, and they motivate work on robustness, interpretability-based monitoring and adversarial training.

As an engineer

If your app's safety depends on the model refusing, assume a determined user will get past it; put the real limits in code, permissions and output checks.

Modern systems that depend on it

  • red-teaming programs
  • input and output safety classifiers
  • adversarial training data
  • agent permission design

Historical context

Before

Safety tuning was evaluated on straightforward harmful requests, and refusals on those were taken as evidence of safety.

After

Safety is evaluated against adaptive attacks, including automated ones, and treated as one layer in a system that limits what any single model output can do.

Used today

Red-team evaluations before releases, attack benchmarks, safety classifiers on inputs and outputs, and permission boundaries in agents.

What to remember

  • Jailbreak: get a safety-trained model to produce what it was trained to refuse. Prompt injection: get an application to follow instructions hidden in data.
  • Competing objectives: helpfulness or instruction-following pulls against refusing (role-play, 'begin your answer with Absolutely! Here is').
  • Mismatched generalization: capabilities reach inputs that safety training never covered (encodings, other languages).
  • Wei et al. 2023: attacks built on these two ideas succeeded on every prompt in a set of unsafe requests against GPT-4 and Claude v1.3.
  • GCG (2023): an automatically optimized 20-token suffix; 88% on harmful strings for Vicuna-7B, transferred to other models.
  • Defence in depth: training + classifiers + least privilege + monitoring. Limit the damage, not just the attempts.

Key papers

Essential

Jailbroken: How Does LLM Safety Training Fail?

Alexander Wei, Nika Haghtalab, Jacob Steinhardt · 2023

Explains why jailbreaks work, in two ideas that still organise the field.

How to read it: Read Section 3 for the two failure modes; the specific attacks have mostly been patched, the reasons have not gone away.

~30 min readarXiv:2307.02483✓ verified 2026-10-07

Watch