AI Feedback: RLAIF and Constitutional Methods

When the annotator is a model: AI-generated preferences, constitutions and critique loops, and the honest limits of self-supervision.

Content last verified 2026-09.

Lessons

  1. Scaling the Annotator
  2. Constitutions and Critique Loops
  3. The Limits of Self-Supervision

Sources