Evaluation
Benchmarks, contamination, judges, and eval harnesses you can defend
- Benchmarks and Their Limits — What a benchmark actually measures, how benchmarks saturate and get gamed, and how to read a leaderboard number without being fooled. (4 lessons, 45 min)
- Data Contamination — When the test set leaks into the training set: why contamination happens at web scale, how it is detected, and what a score means once it has. (3 lessons, 40 min)
- Human Eval, Arenas, and LLM Judges — When there is no answer key: human evaluation, preference arenas and their ratings, and models judging models — with the biases of each. (3 lessons, 45 min)
- Building an Eval Harness — Your benchmark, for your task: golden sets from real traffic, the grader ladder, and a harness you can rerun on every change. (4 lessons, 45 min)
- Production Evals and Regression Testing — Evals as the gate: regression testing before any model or prompt change ships, quality monitoring after, and surviving vendor model updates. (3 lessons, 40 min)
- In Production: Reading Model Evals Critically — The domain’s production capstone: how to read a vendor eval table, why cross-vendor comparisons mislead, and why your own numbers always win. (3 lessons, 40 min)