Concept · Chapter 16: Evaluation, Reliability, Safety & Interpretability
Robustness and Prompt Sensitivity
A model is robust when small changes that should not matter (rephrasing, formatting, a few pixels, a shift in the kind of input) leave its behaviour unchanged, and neural networks often are not.
The problem
Accuracy measured on one format of one test set can collapse under paraphrase, new formatting, a different population or a deliberately crafted input.
The solution
Test with variations (paraphrases, formats, perturbations, shifted data), report the spread rather than one number, search for worst cases, and train or design defences where the stakes need them.
The consequence
Robustness is measured as a range and a worst case, not a single score; deliberately crafted inputs are a security problem, not only a statistics one.
You should understand first
- The Turing Test
- Symbolic AI
- Logic and Rules
- Expert Systems
- Knowledge Representation
- The Knowledge-Acquisition Bottleneck
- From Rules to Learning
- Supervised, Unsupervised and Self-Supervised Learning
- Vectors
- Features, Labels and Tasks
- Loss Functions
- Derivatives and Gradients
- Gradient Descent
- Linear Regression
- Probability and Distributions
- Expected Value and Variance
- Sampling and Uncertainty
- Generalization, Overfitting and Underfitting
- Distribution Shift
- Robustness and Prompt Sensitivity
Changes that should not matter
A person who can answer "What is the capital of France?" can also answer "capital of france??". Models usually can too, but not always, and the failures are often surprising. Three kinds of change are worth separating:
- Harmless variation: paraphrases, typos, formatting, order of options.
- Natural shift: the inputs in use differ from the test set (distribution shift): new products, new slang, another language.
- Adversarial inputs: someone searches for an input that makes the model fail.
Formatting is a variable
Sclar and colleagues sampled many semantically equivalent prompt formats and found accuracy differences of up to 76 points for LLaMA-2-13B, with sensitivity persisting in larger models, with more in-context examples and after instruction tuning. Established HELM made robustness, measured under perturbations of the inputs, one of its seven standard metrics. Established If a result holds only in one format, it is a result about that format.
Adversarial examples
Szegedy and colleagues found that a hardly perceptible perturbation, chosen to maximize a network's error, makes it misclassify an image, and that the same perturbation often fools other networks trained on different data. Established Goodfellow and colleagues attributed this to models behaving too linearly in high dimensions and introduced the fast gradient sign method. EstablishedTiny example. A linear score over 1,000 inputs. Change each input by only in the direction of the sign of its weight. The score moves by ; if the weights average 0.5 in size, that is , enough to flip a confident decision, though no single input changed noticeably. More dimensions, more room for an attacker.
For language models
Text is discrete, so attacks search over tokens rather than nudging pixels, but the logic is the same: use the model's own gradients or feedback to find inputs it handles badly. Zou and colleagues optimized a 20-token suffix that made safety-trained models comply with harmful requests, and it transferred to other models. Established That turns robustness into a security question, which jailbreaks take up.
Measuring robustness
- Run each test item in several formats and paraphrases; report the mean and the spread.
- Keep a slice of the test set from a different time or source to estimate shift.
- For security-relevant uses, report attack success rates under a stated attacker, not average-case accuracy.
Mini experiment
Take five questions your team asks an assistant. Write each three ways: formal, terse with typos, and with the options in a different order. Run all fifteen. How many answers changed? Which version would you have put in a benchmark?
What to remember
- Three kinds of change: harmless variation (paraphrase, format), natural shift (new users, new data), and adversarial inputs (crafted to fail).
- Prompt format alone moved one model's accuracy by up to 76 points (Sclar et al. 2023).
- Adversarial examples (2013): imperceptible image changes that flip a classifier, often transferring between models.
- FGSM: move every input value by ε in the sign of the loss gradient; many tiny changes add up in high dimensions.
- Adversarial suffixes for LLMs are found the same way: optimize the input against the model.
- Report robustness as a spread or worst case, not one number.
Key papers
Holistic Evaluation of Language Models
Percy Liang, Rishi Bommasani et al. · 2022
HELM argued that a single accuracy hides most of what matters and evaluated 30 models on the same scenarios with seven metrics each.
How to read it: The paper is very long; read the introduction's figure of scenarios × metrics and the summary of 25 findings, then dip into a scenario you care about.
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Melanie Sclar, Yejin Choi et al. · 2023
Measured how much a benchmark score can move when only meaningless formatting changes, a hidden source of variance in every comparison.
Universal and Transferable Adversarial Attacks on Aligned Language Models
Andy Zou, Zifan Wang et al. · 2023
Showed that jailbreaks can be found automatically by optimization, like adversarial examples for images, and that they transfer between models.
Intriguing properties of neural networks
Christian Szegedy, Wojciech Zaremba et al. · 2013
Introduced adversarial examples, and in the same paper observed that meaning lives in directions of activation space rather than single units: two threads that run through this chapter.
Explaining and Harnessing Adversarial Examples
Ian J. Goodfellow, Jonathon Shlens, Christian Szegedy · 2014
Explained adversarial examples by linearity, gave the one-step attack everyone learns first, and used it for adversarial training.