Skip to content
Road to Intelligence

Concept · Chapter 16: Evaluation, Reliability, Safety & Interpretability

Robustness and Prompt Sensitivity

Should knowUnderstand12 minDifficulty

A model is robust when small changes that should not matter (rephrasing, formatting, a few pixels, a shift in the kind of input) leave its behaviour unchanged, and neural networks often are not.

The problem

Accuracy measured on one format of one test set can collapse under paraphrase, new formatting, a different population or a deliberately crafted input.

The solution

Test with variations (paraphrases, formats, perturbations, shifted data), report the spread rather than one number, search for worst cases, and train or design defences where the stakes need them.

The consequence

Robustness is measured as a range and a worst case, not a single score; deliberately crafted inputs are a security problem, not only a statistics one.

Changes that should not matter

A person who can answer "What is the capital of France?" can also answer "capital of france??". Models usually can too, but not always, and the failures are often surprising. Three kinds of change are worth separating:

  1. Harmless variation: paraphrases, typos, formatting, order of options.
  2. Natural shift: the inputs in use differ from the test set (distribution shift): new products, new slang, another language.
  3. Adversarial inputs: someone searches for an input that makes the model fail.

Formatting is a variable

Sclar and colleagues sampled many semantically equivalent prompt formats and found accuracy differences of up to 76 points for LLaMA-2-13B, with sensitivity persisting in larger models, with more in-context examples and after instruction tuning. Established HELM made robustness, measured under perturbations of the inputs, one of its seven standard metrics. Established If a result holds only in one format, it is a result about that format.

Adversarial examples

Szegedy and colleagues found that a hardly perceptible perturbation, chosen to maximize a network's error, makes it misclassify an image, and that the same perturbation often fools other networks trained on different data. Established Goodfellow and colleagues attributed this to models behaving too linearly in high dimensions and introduced the fast gradient sign method. Established

Tiny example. A linear score w⋅xw \cdot x over 1,000 inputs. Change each input by only ε=0.01\varepsilon = 0.01 in the direction of the sign of its weight. The score moves by ε∑i∣wi∣\varepsilon \sum_i |w_i|; if the weights average 0.5 in size, that is 0.01×1,000×0.5=50.01 \times 1{,}000 \times 0.5 = 5, enough to flip a confident decision, though no single input changed noticeably. More dimensions, more room for an attacker.

For language models

Text is discrete, so attacks search over tokens rather than nudging pixels, but the logic is the same: use the model's own gradients or feedback to find inputs it handles badly. Zou and colleagues optimized a 20-token suffix that made safety-trained models comply with harmful requests, and it transferred to other models. Established That turns robustness into a security question, which jailbreaks take up.

Measuring robustness

  • Run each test item in several formats and paraphrases; report the mean and the spread.
  • Keep a slice of the test set from a different time or source to estimate shift.
  • For security-relevant uses, report attack success rates under a stated attacker, not average-case accuracy.

Mini experiment

Take five questions your team asks an assistant. Write each three ways: formal, terse with typos, and with the options in a different order. Run all fifteen. How many answers changed? Which version would you have put in a benchmark?

What to remember

  • Three kinds of change: harmless variation (paraphrase, format), natural shift (new users, new data), and adversarial inputs (crafted to fail).
  • Prompt format alone moved one model's accuracy by up to 76 points (Sclar et al. 2023).
  • Adversarial examples (2013): imperceptible image changes that flip a classifier, often transferring between models.
  • FGSM: move every input value by ε in the sign of the loss gradient; many tiny changes add up in high dimensions.
  • Adversarial suffixes for LLMs are found the same way: optimize the input against the model.
  • Report robustness as a spread or worst case, not one number.

Key papers

Important

Holistic Evaluation of Language Models

Percy Liang, Rishi Bommasani et al. · 2022

HELM argued that a single accuracy hides most of what matters and evaluated 30 models on the same scenarios with seven metrics each.

How to read it: The paper is very long; read the introduction's figure of scenarios × metrics and the summary of 25 findings, then dip into a scenario you care about.

~1 h readarXiv:2211.09110✓ verified 2026-10-07
Important

Intriguing properties of neural networks

Christian Szegedy, Wojciech Zaremba et al. · 2013

Introduced adversarial examples, and in the same paper observed that meaning lives in directions of activation space rather than single units: two threads that run through this chapter.

~25 min readarXiv:1312.6199✓ verified 2026-10-07
Optional

Explaining and Harnessing Adversarial Examples

Ian J. Goodfellow, Jonathon Shlens, Christian Szegedy · 2014

Explained adversarial examples by linearity, gave the one-step attack everyone learns first, and used it for adversarial training.

~25 min readarXiv:1412.6572✓ verified 2026-10-07