LLM-as-Judge: An Instrument You Calibrate, Not an Oracle You Trust

A model grading a model is a measuring instrument with known biases and a drift problem. Learn when a judge is the right tool and when it is lazy engineering, how to write rubrics a judge can actually apply, the five biases and their mitigations, and the calibration protocol that turns a judge score into evidence.

Content current as of 2026-09.

Lessons

  1. When a judge is the right instrument — and when it is laziness
  2. Rubric design: broken, then fixed
  3. The five biases, each with its mitigation
  4. Calibration: the number means nothing until you check it
  5. Judging trajectories, and diagnosing disagreement