Skip to content
Road to Intelligence

Concept · Chapter 10: From Base Model to Assistant

Constitutional AI and AI Feedback

Should knowUnderstand10 minDifficulty

Explicit principles can guide model critiques and revisions and help generate preference feedback for later training.

The problem

Human review is expensive, and vague helpfulness or safety judgments are hard to apply consistently.

The solution

Use written principles to guide AI critique, revision and comparisons, then train on the resulting demonstrations or feedback.

The consequence

Some feedback can scale beyond direct human labeling, while depending on human-selected principles and the judging model.

Write the principle before the judgment

A principle might ask whether a response is needlessly harmful or whether it gives a useful benign alternative. A model can use such a principle to critique a draft and write a revision. Those revisions become supervised targets.

The Constitutional AI paper also describes AI comparisons used to train a preference model and a reinforcement-learning stage. The supervised critique/revision stage and the feedback/RL stage are distinct.

What becomes automatic?

Some per-response annotation. The principles, judge choice, prompts and evaluation remain human design decisions. “No human feedback” would obscure that dependence. RLAIF describes the source of feedback; it does not establish that the labels are unbiased or correct.

UltraFeedback is a separate example of scaled AI annotation with multiple judgment dimensions. Do not describe every model-generated label as Constitutional AI: not all use that paper's recipe or an explicit constitution.

A useful disagreement test

Give two judges the same pair under the same principle. Keep disagreements for review. Then alter the wording of the principle and see which decisions change. This is an evaluation proposal, not a claim about a particular released model. A system that merely becomes more consistent with its own judge can still fail an independent test.

What to remember

  • Critique and revision can supply demonstrations; comparisons can supply preference labels.
  • Human-designed principles still shape AI feedback.
  • An automated judge can reproduce or amplify its own mistakes.

Key papers

Important

Constitutional AI: Harmlessness from AI Feedback

Yuntao Bai, Saurav Kadavath et al. · 2022

A documented route for scaling some feedback while retaining human-selected principles.

How to read it: Distinguish the supervised revision stage from the feedback/RL stage.

~30 min readarXiv:2212.08073✓ verified 2026-10-04
Important

UltraFeedback: Boosting Language Models with Scaled AI Feedback

Ganqu Cui, Lifan Yuan et al. · 2023

A concrete source of synthetic feedback used in open assistant research.

How to read it: Distinguish original annotations from a binarized derivative dataset.

~25 min readarXiv:2310.01377✓ verified 2026-10-04

Watch