Concept · Chapter 10: From Base Model to Assistant
Evaluating Post-Training
Evaluate a post-trained assistant on held-out task success, factuality, uncertainty, safety and retained capability, separately from its training reward.
The problem
A falling preference loss or rising reward can conceal hallucination, excessive refusal, memorization or loss of useful skills.
The solution
Keep development and unseen evaluations separate, use independent checks, and compare checkpoints under matched prompts and decoding settings.
The consequence
The result is an evidence-based account of what changed instead of one flattering aggregate score.
You should understand first
- The Turing Test
- Symbolic AI
- Logic and Rules
- Expert Systems
- Knowledge Representation
- The Knowledge-Acquisition Bottleneck
- From Rules to Learning
- Supervised, Unsupervised and Self-Supervised Learning
- Vectors
- Features, Labels and Tasks
- Loss Functions
- Derivatives and Gradients
- Gradient Descent
- Linear Regression
- Probability and Distributions
- Entropy
- Softmax
- Cross-Entropy Loss
- Logistic Regression
- Evaluation Metrics for Classifiers
- Text as Data
- Conditional Probability and Bayes' Theorem
- Probability of Sequences
- Language Modeling
- Dot Product
- Embeddings
- Attention
- Self-Attention
- Causal Masking
- Autoregressive Next-Token Prediction
- Decoding: Greedy, Temperature, Top-k, Top-p
- Learning from Comparisons
- Expected Value and Variance
- Sampling and Uncertainty
- Generalization, Overfitting and Underfitting
- Data Leakage
- Evaluating Post-Training
An assistant needs more than one score
| Question | Example check | What a single preference score can miss |
|---|---|---|
| Did it follow instructions? | Exact sentence count or valid schema | A fluent answer ignores constraints. |
| Is it faithful? | Compare a summary's claims with its source | A polished answer invents a date. |
| Does it know when to stop? | Missing-information prompts | Guessing or unnecessary refusal. |
| Is refusal appropriate? | Problematic requests and benign lookalikes | Blanket refusal looks safe on a one-sided test. |
| Did useful ability regress? | Held-out tasks before and after tuning | A preferred style hides a task regression. |
Keep the judge separate from the exam
Optimizing against one scoring function is not evidence of improvement under another. The reward-overoptimization paper makes that distinction explicit with a synthetic gold reward. Our three-action lab makes the same logical problem visible with invented values; its curves are not a reproduction of the paper's experiments.
Sycophancy is a different failure: agreeing with a user's stated belief when a more accurate answer would disagree. The cited study investigates how feedback and preferences can reward that behavior. It does not establish that every preference-trained model always behaves this way.
A useful comparison protocol
- Freeze a development set and an unseen test set at the prompt/task level.
- Compare the base, SFT and preference-tuned checkpoints with matched generation settings and valid templates for each.
- Blind model identity and randomize answer order for human comparisons.
- Record rubric, ties, disagreement, sample size and uncertainty.
- Use separately checked answers for factual tasks and inspect regressions, not only average wins.
Tulu 3 provides an open case study with development and unseen evaluations. The broader lesson is reusable: selecting a recipe repeatedly on a benchmark makes that benchmark part of development. Keep another test for the final claim.
Try it: in Reward and Restraint, lower β with the flawed judge. Proxy reward increases while correctness falls. Explain which independent measure caught the failure and why KL alone could not.
What to remember
- A training reward and an independent evaluation answer different questions.
- A judge may prefer confidence, verbosity or agreement over accuracy.
- Report prompt splits, rubric, sample size and uncertainty with comparison scores.
Key papers
Scaling Laws for Reward Model Overoptimization
Leo Gao, John Schulman, Jacob Hilton · 2022
Explains why maximizing a learned judge can diverge from the intended outcome.
How to read it: Read the synthetic-gold-reward setup before interpreting the scaling curves.
Towards Understanding Sycophancy in Language Models
Mrinank Sharma, Meg Tong et al. · 2023
Separates being preferred from being accurate.
How to read it: Inspect task construction and judge behavior before generalizing to all feedback-trained models.
Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Nathan Lambert, Jacob Morrison et al. · 2024
Provides an open case study connecting data, objectives and evaluation.
How to read it: Read the evaluation split and training stages; leave reasoning RL detail for Chapter 14.
Watch
Stanford Online
Stanford CS336 Language Modeling from Scratch | Spring 2025 | Lecture 15: Alignment - SFT/RLHF
A technical companion to the chapter’s demonstration, preference and reinforcement-learning pipeline.