Evaluation & Observability

Traces, evals, LLM judges, and knowing whether your agent actually works