Concept · Chapter 10: From Base Model to Assistant
Safety Tuning and Refusal
Safety tuning uses examples and feedback to shape when an assistant should answer, decline or offer a safer alternative.
The problem
Always complying can be harmful, while refusing every sensitive-looking request makes an assistant unusable.
The solution
Train and evaluate on both problematic requests and benign lookalikes, with explicit expectations for appropriate responses.
The consequence
Safety behavior becomes a measured tradeoff rather than a count of refusals, and still needs broader system evaluation.
You should understand first
- Text as Data
- Probability and Distributions
- Conditional Probability and Bayes' Theorem
- Probability of Sequences
- Language Modeling
- Vectors
- Dot Product
- Embeddings
- Attention
- Softmax
- Self-Attention
- Multi-Head Attention
- Causal Masking
- Positional Encoding
- Residual Connections
- Layer Normalization
- Feed-Forward Sublayer (MLP)
- The Transformer Block
- Encoder, Decoder & Encoder–Decoder
- Pretrain, Then Fine-Tune
- Entropy
- Loss Functions
- Cross-Entropy Loss
- Supervised Fine-Tuning
- The Turing Test
- Symbolic AI
- Logic and Rules
- Expert Systems
- Knowledge Representation
- The Knowledge-Acquisition Bottleneck
- From Rules to Learning
- Supervised, Unsupervised and Self-Supervised Learning
- Autoregressive Next-Token Prediction
- Decoding: Greedy, Temperature, Top-k, Top-p
- Learning from Comparisons
- Safety Tuning and Refusal
Two ways to fail
An assistant can comply with a request it should decline. It can also refuse a harmless request because it resembles one seen in safety training. A model that refuses everything would look excellent on a test containing only disallowed requests and terrible on ordinary use.
That is Chapter 3's threshold problem in a richer setting. Evaluate both kinds of error. The response itself matters too: a short explanation and a relevant benign alternative can be more useful than a generic block of boilerplate.
Examples, preferences and principles
SFT can demonstrate suitable answers and refusals. Preference feedback can compare alternatives. Explicit principles can guide model-generated critiques. These are training methods; each still depends on the coverage and quality of the examples or feedback.
The helpful/harmless assistant and Constitutional AI papers provide documented approaches. Their results are bounded by their training and evaluation setups, not proof that safety has been solved.
Keep a paired evaluation
For a teaching evaluation, pair an inappropriate request with a harmless discussion of the same subject. Record compliance quality, unnecessary refusal and the usefulness of redirection separately. Have a rubric before looking at model outputs.
What to remember
- Refusal rate alone is not safety quality.
- Test unsafe compliance and unnecessary refusal separately.
- Post-training addresses parts of alignment, not the entire system problem.
Key papers
Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Yuntao Bai, Andy Jones et al. · 2022
Makes feedback data and the helpfulness/harmlessness tension concrete.
How to read it: Inspect the data and evaluation categories; do not reduce safety to refusal rate.
Constitutional AI: Harmlessness from AI Feedback
Yuntao Bai, Saurav Kadavath et al. · 2022
A documented route for scaling some feedback while retaining human-selected principles.
How to read it: Distinguish the supervised revision stage from the feedback/RL stage.
Watch
Stanford Online
Stanford CS336 Language Modeling from Scratch | Spring 2025 | Lecture 15: Alignment - SFT/RLHF
A technical companion to the chapter’s demonstration, preference and reinforcement-learning pipeline.