DPO and the Direct Methods
Direct preference optimization skips the reward model entirely. The trick, the math, the variants — and when classic RLHF still wins.
Content last verified 2026-09.
Lessons
Sources
- Rafailov et al. (2023) — Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- Ouyang et al. (2022) — Training language models to follow instructions with human feedback (InstructGPT)
- Gao, Schulman & Hilton (2022) — Scaling Laws for Reward Model Overoptimization
- Azar et al. (2023) — A General Theoretical Paradigm to Understand Learning from Human Preferences (IPO)