Eval-authoring agents.
Everybody who has shipped an agent has the same bottleneck — a hand-written eval set of forty cases that stopped discriminating months ago — and the obvious fix, pointing a model at it, fails for one specific reason: the thing that makes test-generation agents work is a free oracle, and agent evaluation does not have one. Build this agent and the only division of labour that survives contact with production is narrow: the agent generates situations, and something that is not the agent owns the answer.
Name the oracle before you write a line of generator.
A unit-test generator gets its correctness signal for free. The compiler types the call, the runner reports pass or fail, and for regression work the current behaviour is the specification — golden-master capture turns "what should this return?" into "what does this return today?" with no judgement required. That is why test generation works at scale and why its metrics look so good.
Now ask the same question of an agent eval: did the refund agent handle this customer correctly? There is no runner. The entity best equipped to answer is a model of roughly the capability you are trying to measure, which is the generator–verifier gap arriving as a staffing problem. So fix the vocabulary before you build anything:
- A situation is the inputs: user goal, conversation prefix, tool and data state, and any injected adversity. Cheap to generate, cheap to check for plausibility, and genuinely under-supplied — this is what you want the agent producing by the thousand.
- An oracle is the thing that decides whether a trajectory through that situation was correct. Expensive, contested, and the only part that carries authority. This is the part the agent does not own.
- An eval case is a situation bound to an oracle. Count these, not situations, when you report progress — a thousand situations with no oracle is a load test, not an eval set.
The one-sentence rule that prevents the expensive mistake: if the same model family writes both the case and the grader, you have built a system that optimises the grader. It will report improvement indefinitely. Everything in the rest of this playbook is machinery for keeping those two jobs apart.
Mine traces for shapes; generate for coverage of those shapes.
A generator asked for "hard customer-support cases" produces a plausible, narrow, synthetic-smelling set that your agent passes. The fix is to stop treating generation as invention and treat it as sampling from a distribution you already own: your production traces. Two passes, in this order.
Pass one — cluster real traffic into failure shapes. Pull the runs that went wrong (thumbs-down, escalation, retry storms, abandoned sessions), cluster them by embedding, and have the agent write one sentence naming each cluster. You are after names like "user supplied an order number that belongs to a different account" or "the tool returned data from before the user's last edit", not severity scores. Ten to thirty named shapes is the realistic yield from a few thousand traces, and that list is the actual deliverable of pass one.
Pass two — generate along named axes within each shape. Now the generator has somewhere to go. Give it the shape, an example trace, and a fixed axis list to perturb along:
- Missing or malformed input — a required field absent, a date in the wrong locale, an identifier that is syntactically valid and refers to nothing.
- Two defensible readings — a request where a competent human would ask a clarifying question. These are where agents silently pick one and proceed.
- Mid-task change of goal — the user contradicts an earlier instruction on turn four. Tests whether the agent re-plans or keeps executing the dead plan.
- Stale or inconsistent tool state — the read returns a value the write already superseded.
- Adversarial content in retrieved data — instructions embedded in a document the agent was asked to summarise.
Deduplicate aggressively — embed every candidate situation and drop anything within a tight cosine threshold of an accepted case — and auto-reject any candidate whose correct outcome cannot be stated in one sentence. That rejection rule alone removes most of what a generator produces, and everything it removes would have cost you a human reviewer's hour.
Three grader classes, and what the agent may touch in each.
Rank every case by the strongest oracle it can get, and spend your human attention on the cases that cannot climb:
- Programmatic state assertion — the refund exists in the ledger with amount X; the ticket is in state
resolved; the file parses and contains three rows. Deterministic, cheap to re-run, immune to judge drift. The agent may draft these assertions, and a human reviews the diff, exactly as you would review generated test code. Target getting as much of your set as possible into this class — it is worth rewriting a case to make it state-assertable. - Rubric plus calibrated judge — for cases where correctness is a quality, not a state. The agent may propose the rubric, never the answer key, and never both for the same case. A judge is admissible only once you have measured its agreement against human labels on a held-out slice; see judge calibration and LLM-as-judge for agents.
- Human only — the residue. Keep it small and keep it, because it is your calibration anchor; the day you have no human-labelled slice is the day judge drift becomes invisible. Annotation ops is the machinery for running it.
One hard constraint, cheap to enforce in code and expensive to discover later: the model under test must never be the judge of its own family, and the generator's model must never be the judge at all. Record generator model, judge model and system-under-test model on every case, and fail the eval run — loudly, not with a warning — when any two of them match.
The highest-value cases are the ones with no right answer.
Nearly every hand-written eval set is 100% satisfiable, which means it cannot detect the production failure that costs the most: an agent that cannot do the job and reports that it did. That is failure concealment, and it is invisible to a set in which every task is completable.
So brief the generator explicitly for unsatisfiable work, and make it a quota rather than a nice-to-have. Roughly a fifth to a quarter of a mature set should be cases where the honest answer is a refusal or a question:
- Permission absent — the task requires a scope the agent's credential does not carry.
- Referent absent — the record, file or account named in the request does not exist.
- Internally contradictory — two constraints in the request cannot both hold.
- Outside policy — completable, but the agent should decline.
- Underdetermined — missing a fact only the user has; the correct action is one clarifying question, and a confident answer is a failure even when it happens to be right.
Score these the same way you score the rest, with full credit for a correct refusal and zero for a fabricated completion — and never label them in anything the system under test can see. A generator that is good at this is more valuable than a generator that is good at anything else, because this is the class your humans systematically fail to write.
Provenance per case, or the set rots silently.
A generated set is a dataset with a supply chain, and the failure mode is not a broken case — it is a set that still runs, still reports numbers, and no longer measures anything. Carry metadata on every case from the moment it is created:
case_id: rfnd-0412
situation: { goal, conversation_prefix, tool_state, injected }
oracle: { class: state_assertion | rubric_judge | human, ref }
satisfiable: false
shape: "order id belongs to another account"
source_trace: trace_8e21c4 (2026-09-12)
generator: model=<id> prompt_version=7
reviewer: alice (2026-09-14)
frozen: true # excluded from any prompt or training corpus
discriminates: [v3.1 > v2.8] # last pair of systems this case separated
Four policies hang off those fields:
- Freeze a private slice and never ship it anywhere. Generated cases are text, and text leaks into prompts, fine-tuning corpora and few-shot examples. Contamination is not a public-benchmark problem; it is what happens when someone pastes a failing eval case into a prompt to debug it.
- Regenerate on drift, not on a calendar. The trigger is your trace clusters moving — a new shape appearing, or an old one ceasing to occur. Monthly regeneration on a quiet product is churn; see eval-set maintenance.
- Version the set and pin runs to a version. A score is meaningless without one, and a set that grows between two runs makes the comparison unpaired — which, per eval variance and statistical power, is where most phantom improvements come from.
- Retire cases that stop discriminating. A case every system passes costs money and tells you nothing. Move it to a cheap smoke set and take it out of the headline number.
The gate: does the generated set separate systems you already know differ?
The eval-authoring agent needs its own acceptance test, and it is not "how many cases did it write". Take two versions of your agent whose relative quality you are confident about — ideally one you deliberately degraded — and ask whether the new cases rank them correctly. A case that both versions pass, or both fail, has no discriminating power regardless of how clever it looks.
- Discrimination rate — share of accepted cases that separate a known-better system from a known-worse one. This is the primary metric, and it is the one nobody reports.
- Verdict stability — run each case three times against an unchanged system; a case whose verdict flips is measuring sampling noise. Report single-attempt success alongside all-attempts success: the gap is large and widens with repeats, as on τ²-Bench, where one reported configuration's all-of-k success fell from 81.6% at k=1 to 56.1% at k=4 on the retail domain.
- Human minutes per accepted case — the real unit cost. If review takes longer than writing the case by hand would have, the generator is net negative and the fix is upstream, in the rejection rules of STEP 2.
- Acceptance rate — accepted over generated. Below roughly one in five, stop generating and go fix the shape list; the generator is not short of fluency, it is short of grounding.
Start here, in this order, and you will have something useful in a week: cluster last month's bad runs into named shapes by hand; pick the three shapes that hurt most; have the agent generate twenty situations per shape with a hard "state the correct outcome in one sentence" filter; convert everything you can into state assertions and review them like code; and require that a fifth of the set be unsatisfiable with credit for refusing. Then hold the whole thing to one bar — a case earns its place only by separating two systems you already know differ. Everything else in this playbook is an elaboration of that bar.
Related: simulated users for generating the conversation side of a situation, eval integrity and scorer gaming for what happens when the grader becomes the target, and eval-driven development for wiring the resulting set into a gate.