RL for Reasoning Models
Verifiable rewards changed the game: training models to think longer on math and code, what reasoning training buys, and what it costs at inference time.
Content last verified 2026-09.
Lessons
Sources
- DeepSeek-AI (2025) — DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Gao, Schulman & Hilton (2022) — Scaling Laws for Reward Model Overoptimization
- Ouyang et al. (2022) — Training language models to follow instructions with human feedback (InstructGPT)
- Schulman et al. (2017) — Proximal Policy Optimization Algorithms