Skip to content
Road to Intelligence

Concept · Chapter 10: From Base Model to Assistant

Evaluating Post-Training

Must knowKnow well12 minDifficulty

Evaluate a post-trained assistant on held-out task success, factuality, uncertainty, safety and retained capability, separately from its training reward.

The problem

A falling preference loss or rising reward can conceal hallucination, excessive refusal, memorization or loss of useful skills.

The solution

Keep development and unseen evaluations separate, use independent checks, and compare checkpoints under matched prompts and decoding settings.

The consequence

The result is an evidence-based account of what changed instead of one flattering aggregate score.

An assistant needs more than one score

QuestionExample checkWhat a single preference score can miss
Did it follow instructions?Exact sentence count or valid schemaA fluent answer ignores constraints.
Is it faithful?Compare a summary's claims with its sourceA polished answer invents a date.
Does it know when to stop?Missing-information promptsGuessing or unnecessary refusal.
Is refusal appropriate?Problematic requests and benign lookalikesBlanket refusal looks safe on a one-sided test.
Did useful ability regress?Held-out tasks before and after tuningA preferred style hides a task regression.

Keep the judge separate from the exam

Optimizing against one scoring function is not evidence of improvement under another. The reward-overoptimization paper makes that distinction explicit with a synthetic gold reward. Our three-action lab makes the same logical problem visible with invented values; its curves are not a reproduction of the paper's experiments.

Sycophancy is a different failure: agreeing with a user's stated belief when a more accurate answer would disagree. The cited study investigates how feedback and preferences can reward that behavior. It does not establish that every preference-trained model always behaves this way.

A useful comparison protocol

  1. Freeze a development set and an unseen test set at the prompt/task level.
  2. Compare the base, SFT and preference-tuned checkpoints with matched generation settings and valid templates for each.
  3. Blind model identity and randomize answer order for human comparisons.
  4. Record rubric, ties, disagreement, sample size and uncertainty.
  5. Use separately checked answers for factual tasks and inspect regressions, not only average wins.

Tulu 3 provides an open case study with development and unseen evaluations. The broader lesson is reusable: selecting a recipe repeatedly on a benchmark makes that benchmark part of development. Keep another test for the final claim.

Try it: in Reward and Restraint, lower β with the flawed judge. Proxy reward increases while correctness falls. Explain which independent measure caught the failure and why KL alone could not.

What to remember

  • A training reward and an independent evaluation answer different questions.
  • A judge may prefer confidence, verbosity or agreement over accuracy.
  • Report prompt splits, rubric, sample size and uncertainty with comparison scores.

Key papers

Important

Scaling Laws for Reward Model Overoptimization

Leo Gao, John Schulman, Jacob Hilton · 2022

Explains why maximizing a learned judge can diverge from the intended outcome.

How to read it: Read the synthetic-gold-reward setup before interpreting the scaling curves.

~30 min readarXiv:2210.10760✓ verified 2026-10-04
Important

Towards Understanding Sycophancy in Language Models

Mrinank Sharma, Meg Tong et al. · 2023

Separates being preferred from being accurate.

How to read it: Inspect task construction and judge behavior before generalizing to all feedback-trained models.

~25 min readarXiv:2310.13548✓ verified 2026-10-04
Important

Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Nathan Lambert, Jacob Morrison et al. · 2024

Provides an open case study connecting data, objectives and evaluation.

How to read it: Read the evaluation split and training stages; leave reasoning RL detail for Chapter 14.

~40 min readarXiv:2411.15124✓ verified 2026-10-04

Watch