Human Eval, Arenas, and LLM Judges

When there is no answer key: human evaluation, preference arenas and their ratings, and models judging models — with the biases of each.

Content last verified 2026-09.

Lessons

  1. Human Evaluation
  2. Arenas and Ratings
  3. LLM as Judge

Sources