RL for Reasoning Models

Verifiable rewards changed the game: training models to think longer on math and code, what reasoning training buys, and what it costs at inference time.

Content last verified 2026-09.

Lessons

  1. Verifiable Rewards
  2. Training Models to Think Longer
  3. What It Buys — and What It Costs

Sources