Operations / Evaluation & Observability
Evaluation & Observability
Measuring agents that don't have a single right answer — outcome vs trajectory evals, LLM-as-judge, traces, benchmarks.
- Why Evaluating Agents Is HardNon-determinism, compounding multi-step error, no single gold answer, path-dependence, eval cost, and dataset rot — the six reasons one clean number is a lie.
- Online vs offline evalsOffline evals catch regressions before deploy; online evals catch the user behavior you couldn't fake — why you need both, and where each one lies to you.
- Outcome vs Trajectory EvaluationEnd-state predicates vs grading the decision sequence: when each is right, partial credit, and tool-call assertions as the highest-leverage safety check.
- LLM-as-Judge for AgentsRubric design, pairwise vs pointwise, the biases that invert verdicts, calibrating against human labels, and the cases where you must not use a judge.
- Reading Agent Benchmarks CriticallyWhat SWE-bench, GAIA, τ-bench and WebArena actually measure, why contamination and harness sensitivity make rank a weak signal, and the small custom set that really decides.
- Tracing & Observability for AgentsThe trace is the data structure, not a log: what to record per step, spans and OpenTelemetry GenAI conventions, and trajectory replay as the bridge to eval.
- Eval-Driven Agent DevelopmentThe eval is the only spec an agent has: tiered CI gates, golden trajectories, offline vs online, the production-to-eval flywheel, and the no-regression ratchet.
- OpenTelemetry GenAI Semantic ConventionsInstrumentation is a data-model decision, not a dashboard one: the agent/workflow/tool/model span kinds, why Development status argues for pinning rather than waiting, splitting structural telemetry from prompt content at the collector, and the one collector hop that makes every later vendor choice reversible.
- Detecting Quality RegressionsProduction has no labels and a judged metric needs ~1,400 scored runs to see a five-point drop, so the detector is the shape of the run — step-cap rate, per-tool errors, termination mix — and the judge is only the confirmation.
- Production Feedback SignalsThumbs arrive from under one percent of sessions and reward confidence over correctness, while the diff between the agent's output and what the user actually shipped is a dense, free, expert-written label — treat every feedback signal as a router into the eval set rather than a metric to optimise.
- Annotation & Labeling OpsA judge cannot be more accurate than the labels it was calibrated against, so if two of your experts agree on 72% of traces a judge at 72% is already at the ceiling — measure inter-annotator agreement first, read low agreement as a rubric defect, and route disagreement to adjudication instead of averaging it away.
- Trace Sampling & RetentionSampling 10% at run start keeps 10% of your failures, and nothing at step zero predicts which run goes wrong — so buffer to run end, keep every failed, capped and expensive run whole, cut successes hard, and treat retention as a decision about the eval set you have not built yet.
- Simulated Users in Agent EvaluationEvery multi-turn agent score measures two systems, and the second one is an unversioned model playing a customer that can move your number several points on its own — pin its model ID, prompt and seed like a dependency, calibrate it against real transcripts, never let it judge whether it was satisfied, and give graders an explicit simulator-fault verdict.
- Measuring Agent LatencyA fifteen-step trajectory turns a one-in-a-hundred slow call into a one-in-seven slow task, so the p99 of a step predicts your users' experience far better than its p50 — measure the trajectory rather than the call, split model from tool from queue time, and keep time-to-first-useful-output separate from time-to-done.
- Maintaining an Eval SetAn eval set decays by being fitted, not by rotting: every regression you fix converts a discriminating case into a permanent pass, so measure the fraction of cases all candidates already pass, score cases by how often they changed anyone's mind, and run a standing replacement rate instead of a periodic cleanup.
- Online Experiments for AgentsDetecting a three-point lift in task success takes about 3,700 sessions per arm before anything agent-specific, and clustering by user typically triples it — so randomise the user rather than the request, pre-commit to one decision metric chosen for its variance, and when the arithmetic says the test is unaffordable, run a guarded rollout and label it as one.
- Failure Taxonomies & Triage"Hallucination" names the smoke at the end of a cascade and routes the ticket to the wrong team — label the earliest step where a competent operator would have acted differently, derive the classes bottom-up from a hundred read traces, weight them back to the prevalence in a uniform random sample, and give every class an owner, a regression case and a detector.
- Redacting PII from Agent TracesContent capture is opt-in for a reason: redact in the SDK before the span leaves the process, emit deterministic typed tokens rather than masks so joins and erasure still work, measure per-entity recall as a CI gate, and size a break-glass raw tier above your measured MTTD.
- Evaluating Guardrails & DetectorsRecall on a balanced benchmark is the number that does not transfer: at a 1-in-10,000 attack base rate a 99%-recall, 1%-false-positive detector yields under 1% precision, so measure your own prevalence from a blind uniform sample, derive the threshold from the cost ratio of the two errors, keep the frozen regression set separate from a red-team set that rotates, and record what the detector adds to p95 and what it does when it times out.
- Shadow Mode & Dark LaunchesA shadow agent never lives with its own mistakes, so its errors do not compound and the observed per-task rate drifts toward the per-step rate — an upper bound, biased worst on exactly the long runs you wanted reassurance about. Write down which effects are suppressed and what the permitted ones cost, mirror a sample rather than all traffic, adjudicate only the disagreements blind in three buckets, and promote on thresholds written before the run.
- Replay Testing with Recorded TracesEvals measure the model and unit tests measure the code; neither notices that a refactor moved authentication to step four. Replay catches that class cheaply, but a recording pins the world, so it proves nothing past the first action taken differently — which makes the cache-miss policy the whole design. Split it into pinned trajectories that assert tool sequence and step count pre-merge, and stateful contract fixtures scored on outcome nightly; stamp every recording and expire it, because a fully green suite replaying a world that stopped existing is how third-party drift reaches production.
- Screenshots & DOM Artefacts in Agent TracesThe first browser agent in production turns your trace store into an image archive, and every control you built assumes text: a redactor at 99% recall on prompts scores zero on a PNG, and a frame is at once the model's input, your only audit evidence, and an unscanned injection channel. Capture the accessibility tree as the text-of-record, mask inside the page before the pixels are ever encoded, keep full frames only at the step before a write and at the last step of a failure, and give blobs their own store and their own clock.
- Refusal Monitoring in ProductionA refusal returns 200, costs fewer tokens and raises no error, so a quality regression presents as a cost win — and in a loop it is a step that silently vanished from a run that reported success: classify four causes, freeze a canary set, and alarm on the delta rather than the level.
- Scope-Conformance EvaluationYour suite answers whether the task finished and has no opinion on what else the agent touched, which is now the axis a frontier lab gated a release on. Vary three things you currently ship unreviewed — the scope clause, what your harness returns when the agent asks a human and none is there, and whether the sanctioned route works — then count actions against targets nobody named, stage by stage against a declared target ledger. The published run of that grid moved full out-of-scope attacks from 26 of 50 trajectories to 4 of 49 on one sentence, and found the agent treating its own harness's filler reply as authorisation in 44% of hard cases. Report the scaffold with the number, and the residual as well as the improvement.