Skip to content
Road to Intelligence

Concept · Chapter 10: From Base Model to Assistant

Safety Tuning and Refusal

Must knowUnderstand9 minDifficulty

Safety tuning uses examples and feedback to shape when an assistant should answer, decline or offer a safer alternative.

The problem

Always complying can be harmful, while refusing every sensitive-looking request makes an assistant unusable.

The solution

Train and evaluate on both problematic requests and benign lookalikes, with explicit expectations for appropriate responses.

The consequence

Safety behavior becomes a measured tradeoff rather than a count of refusals, and still needs broader system evaluation.

Two ways to fail

An assistant can comply with a request it should decline. It can also refuse a harmless request because it resembles one seen in safety training. A model that refuses everything would look excellent on a test containing only disallowed requests and terrible on ordinary use.

That is Chapter 3's threshold problem in a richer setting. Evaluate both kinds of error. The response itself matters too: a short explanation and a relevant benign alternative can be more useful than a generic block of boilerplate.

Examples, preferences and principles

SFT can demonstrate suitable answers and refusals. Preference feedback can compare alternatives. Explicit principles can guide model-generated critiques. These are training methods; each still depends on the coverage and quality of the examples or feedback.

The helpful/harmless assistant and Constitutional AI papers provide documented approaches. Their results are bounded by their training and evaluation setups, not proof that safety has been solved.

Keep a paired evaluation

For a teaching evaluation, pair an inappropriate request with a harmless discussion of the same subject. Record compliance quality, unnecessary refusal and the usefulness of redirection separately. Have a rubric before looking at model outputs.

What to remember

  • Refusal rate alone is not safety quality.
  • Test unsafe compliance and unnecessary refusal separately.
  • Post-training addresses parts of alignment, not the entire system problem.

Key papers

Important

Constitutional AI: Harmlessness from AI Feedback

Yuntao Bai, Saurav Kadavath et al. · 2022

A documented route for scaling some feedback while retaining human-selected principles.

How to read it: Distinguish the supervised revision stage from the feedback/RL stage.

~30 min readarXiv:2212.08073✓ verified 2026-10-04

Watch