Deep-Dives / Evaluating Agents
Evaluating Agents
The 2026 discipline of evaluating agents: benchmark saturation, judge calibration, drift detection, and the eval methodologies that survived contact with production.
- Judge Calibration & Meta-EvaluationPrometheus 2, JudgeBench, RubricEval; meta-evaluation collapse; the 85-90% human-agreement floor; monthly recalibration cadence.
- Benchmark Landscape (2026)SWE-bench Verified saturation (five models within 0.7 pts); SWE-bench Pro; contamination as legal deterrent; why Verified is now an audit signal, not a ranking.
- HAL & Asynchronous Agent EvalPrinceton HAL (cost-per-solve + 5-dim reliability dashboard); Gaia2 (async environments, write-action verifiers, temporal constraints); why static benchmarks miss real deployment.
- Trajectory & Process EvaluationScoring how the agent worked, not just the final answer: outcome vs trajectory eval; AgentEvals match modes (strict/unordered/subset/superset); reference-based vs reference-free LLM-judge; the process-reward-model crossover; and why exact-match on the path fails correct agents.
- Eval-Driven Development & Regression Evals in CIEvals as CI gates, not one-offs: the golden set as a living asset; pass^k and paired significance tests for non-determinism; online vs offline, canary + drift detection; cost/latency as gate-able budgets; promptfoo, DeepEval, Inspect.
- Eval Variance & Statistical PowerA single-run agent score is a sample, not a measurement: pass@k vs pass^k, why between-task variance means more tasks beats more runs, paired McNemar designs that halve the detectable effect on the same budget, and the decision rule that stops teams ratcheting on noise.
- Benchmark Contamination & LeakageContamination belongs to the (model, benchmark, date) triple, not to the benchmark: verbatim vs solution vs indirect leakage, why agent benchmarks leak the environment and not just the answer, four detection tests you can run from outside, and the rule that public scores screen while only post-cutoff data decides.
- Eval Integrity & Scorer GamingA score is produced by software the agent under test can reach: task-level reward hacking, harness compromise and cross-run contamination are three different failures wearing one word, and the third silently fakes the sample rather than one run. Grade out of process from an immutable transcript, delete every shared writable surface between runs, and publish the isolation properties next to the number.
- Evaluating Against Live SystemsAn eval that transacts with systems you do not own is a deployment with no change control, and the September 2026 Medicare-portal incident happened inside one: score a circumvented refusal as a failure or your successful trajectories become a reward for circumvention; then cap per-host load, run from an attributable egress, and keep request logs joined to trajectories with a named owner who may notify.
- Human Baselines in Agent EvalsA score has no denominator until you know what a competent person scores on the same tasks, under the same cap, with the same tools, graded by the same grader — and that number was never collected for almost any benchmark teams quote. A baseline is that four-part tuple, the grader biases the comparison in both directions at once, and METR's time-horizon method is the field's only calibrated answer. Then build your own in two person-days and report the ratio, never the score.