Eval Fundamentals: You Cannot Improve What You Cannot Measure

The centrepiece discipline of agent engineering: turning “it worked when I tried it” into a pass rate on a golden dataset. Build the dataset, separate outcome evals from trajectory evals, climb the grading ladder from deterministic checks to judges, and turn every incident into a permanent test case.

Content current as of 2026-09.

Lessons

  1. Why “it worked when I tried it” is not evidence
  2. The golden dataset: twenty real cases beat a thousand invented ones
  3. Outcome evals and trajectory evals
  4. The grading ladder: deterministic first, judges last
  5. Pass rates and honest variance
  6. The eval lifecycle: every incident becomes a case