RLHF: Learning from Preferences
Reward models trained on human comparisons, policy optimization against them, and the failure mode that haunts the whole method: reward hacking.
Content last verified 2026-09.
Lessons
- Why Preferences, Not Demonstrations
- Reward Models
- The Optimization Loop
- Reward Hacking and Overoptimization
Sources
- Christiano et al. (2017) — Deep Reinforcement Learning from Human Preferences
- Stiennon et al. (2020) — Learning to Summarize from Human Feedback
- Ouyang et al. (2022) — Training Language Models to Follow Instructions with Human Feedback (InstructGPT)
- Schulman et al. (2017) — Proximal Policy Optimization Algorithms
- Gao, Schulman & Hilton (2022) — Scaling Laws for Reward Model Overoptimization