Eval Fundamentals: You Cannot Improve What You Cannot Measure
The centrepiece discipline of agent engineering: turning “it worked when I tried it” into a pass rate on a golden dataset. Build the dataset, separate outcome evals from trajectory evals, climb the grading ladder from deterministic checks to judges, and turn every incident into a permanent test case.
Content current as of 2026-09.