DPO and the Direct Methods

Direct preference optimization skips the reward model entirely. The trick, the math, the variants — and when classic RLHF still wins.

Content last verified 2026-09.

Lessons

  1. Skipping the Reward Model
  2. The DPO Objective
  3. Variants and Trade-offs

Sources