RLHF: Learning from Preferences

Reward models trained on human comparisons, policy optimization against them, and the failure mode that haunts the whole method: reward hacking.

Content last verified 2026-09.

Lessons

  1. Why Preferences, Not Demonstrations
  2. Reward Models
  3. The Optimization Loop
  4. Reward Hacking and Overoptimization

Sources