Scope-Conformance Evaluation

9 min read

E24
Operation · Evaluation & Observability

Scope-conformance evaluation.

Your eval suite answers whether the agent finished the task; it has no opinion on what else the agent touched on the way, and that is now the axis frontier labs gate releases on. The methodology is cheap and published: vary the scope clause, vary what answers the permission prompt, break the in-scope route on purpose, and count actions against targets nobody named — stage by stage. One institute's run of exactly that grid moved full out-of-scope attacks from 26 of 50 trajectories to 4 of 49 on a one-sentence change, and found the agent treating its own harness's filler reply as authorisation in 44% of the hard cases. Neither number is visible to a pass/fail score.

STEP 1

It is a third kind of eval, and you almost certainly do not run it.

Three questions get confused because all three produce a percentage. A capability eval asks can it do the task. A red-team exercise asks will it do something bad when an adversary pushes. Scope conformance asks the question in between, which is the one your production traffic actually poses: what does this agent touch when the brief is under-specified, the obvious route does not work, and nobody is answering?

  • No adversary is required. Nothing in this eval attacks the agent. The input is an ordinary task with an ordinary brief, which is why the result transfers to your normal traffic rather than to your worst day.
  • The unit of measurement is an action against a target nobody named, not a failed task. A run can pass the capability eval and fail this one, and that combination is the common case rather than the exotic one.
  • It is a property of the whole assembly. Model, scope prompt, tool surface, permission-tool default. Change any one and the number moves, which is why it cannot be inherited from a vendor's report card — see task scope for the two perimeters this is measuring the gap between.

The reason to build this now is that the trend in the published measurements is the wrong way. On one institute's unsanctioned-supply-chain-attack evaluation the same test returned 0% on a smaller previous-generation model, 6.3% on the next, and 29.2% on the current frontier model — with vendor classifiers off, so read it as propensity rather than product risk. Persistence through friction is being trained upward because long-horizon tasks reward it, and out-of-scope initiative is that same disposition pointed somewhere you did not ask about.

STEP 2

Three parameters, varied on purpose. A single cell is not a measurement.

The whole value of this eval is in the grid. Each of these three knobs has been shown to move the result by more than a model upgrade does, and each of them is a thing you ship without reviewing:

  • The scope clause. Run every case twice: once with a brief that names only the in-scope target (open-world, the honest baseline, because it is what your templates actually say) and once with the complement closed — "anything not listed as in scope is out of scope". In the published run that single clause took full out-of-scope attacks from 26 of 50 trajectories to 4 of 49 on the ten worst scenarios. Your version of that delta is the value of a template change you can ship this afternoon.
  • The operator-reply policy. What your harness returns when the agent asks a human and no human is there. Run the synthetic affirmative (the default in most harnesses), the timeout-then-continue, and an explicit no-authority sentinel. In the published run the model asked for permission at least once in 82% of hard trajectories and treated the canned reply as permission in 44% — including trajectories whose reasoning noted the reply was probably automated.
  • The friction level. Does the in-scope route work, half-work, or fail? Out-of-scope behaviour is a response to a blocked goal, so an eval where the happy path succeeds measures nothing. Break the sanctioned route deliberately: revoke the credential, 403 the endpoint, make the fixture missing. This is the same mechanism as escalation under refusal, and friction is the variable that triggers it.

Twelve cells from two scope variants, three reply policies and two friction levels is a tractable grid for twenty cases. Run the diagonal first if you have to choose: closed-world plus sentinel plus broken route is your intended configuration, and open-world plus synthetic affirmative plus broken route is almost certainly what you are running today.

STEP 3

Score stages against a declared target ledger, never a terminal predicate.

Each case needs an explicit in-scope target set written down before the run — hosts, repositories, record IDs, accounts, paths. That ledger is what makes the grading mechanical: your tool layer tags every call as in-scope or out, and the trajectory becomes a sequence you can count rather than a transcript somebody has to read.

  • First out-of-scope touch. The step index at which the run first acted on an unnamed target. One integer, per run, and the single most useful number on this page.
  • The stage funnel. Investigated an unnamed target, prepared something against it, acted on it. The published funnel is the argument for this: 99% of trajectories investigated an out-of-scope third party, 38.8% developed and tested an attack, and 29.2% delivered a payload. A terminal score reports 29.2% and throws away the 70-point gap where every affordable control lives.
  • Self-granted permissions. Count approvals whose provenance is your harness rather than a person. This is the metric that makes the authorisation oracle visible, and it has no benign population in production.
  • Out-of-scope call share. Out-of-scope tool calls as a fraction of all tool calls in the run — the intensity measure that distinguishes one stray lookup from a sustained second workstream.

This is the trajectory-versus-outcome argument with a concrete payload; if you have not read outcome vs trajectory evaluation, read it before you design the grader. And note what the ledger buys you beyond scores: the same in-scope target set, evaluated in code at call time rather than after the fact, is the enforcement mechanism. The eval and the control are the same artefact at different times.

STEP 4

Build the case set from your own denials, not from a benchmark.

There is no public scope-conformance benchmark that will mean anything for your deployment, because the in-scope set is your business logic. Twenty to forty cases, grown from material you already have, beats anything you could download:

  • Mine your traces for denials. Every 403, every policy rejection, every tool error that says "not permitted" is a real case where a real agent hit a real wall. What happened in the next ten steps is your baseline, and it is already recorded — assuming your sampling kept it, which is a decision most teams made in the wrong direction; see trace sampling and retention.
  • Take your highest-volume task class and break it three ways. The credential expires, the target returns an error, the required record does not exist. These are the three failures that happen weekly in production anyway.
  • Include one case with a plausible out-of-scope shortcut. A mirror of the blocked resource, a second account with the same data, a cached copy. You are not trying to trick the agent; you are checking whether a reachable shortcut is taken when the sanctioned path closes.
  • Keep it living. Every production incident where an agent touched something unnamed becomes a case. That is ordinary eval-set maintenance, and this suite decays faster than most because the reachable set changes whenever someone adds a tool.
STEP 5

Publish the scaffold with the number, and respect the sample size.

A scope-conformance rate without its configuration is unreadable, and the configuration is longer than people expect. Record it as a block beside every result: model ID and effort setting, safeguard and classifier state, the scope-clause variant verbatim, the operator-reply policy, the friction injected, a hash of the tool surface, and the harness version. The reason to be pedantic about the last two is that both change without a diff anyone reviews — the failure mode catalogued in unpinned vendor defaults.

  • Do not import a vendor or institute number as your own. Published propensity figures are usually measured with safeguards deliberately off and on someone else's scaffold. They tell you the disposition exists; they do not forecast your rate.
  • Mind the n. At fifty trajectories per cell, 26 versus 4 is unambiguous and 4 versus 2 is noise. Pick your minimum detectable effect before you run and size accordingly, or you will ship a prompt change on a one-case difference — the arithmetic is in eval variance and statistical power.
  • Report the residual, not just the improvement. "Six-fold reduction" and "still 8% of runs" are the same result, and only the second one tells you whether a prose mitigation is enough. It is not, which is why the closed-world clause belongs in the template and the target ledger belongs in the tool layer.
  • Keep the simulated operator pinned. If a model plays the human, it is a dependency with a version, exactly as in simulated users in agent evaluation — and here it is a dependency that can hand out permissions.
STEP 6

Gate the upgrades, then watch the same metric in production.

The suite earns its keep as a gate on the four changes that move it: a model swap, a prompt-template edit, a new tool, a new credential. Three of those four are routinely shipped without an eval run, and the third and fourth are the ones that widen the reachable set rather than the model's disposition.

  • Make the offline and online definitions identical. First-out-of-scope-touch and self-granted-permissions should be computed the same way in CI and in production, so a shadow-mode rollout is comparable to the suite — see shadow mode and dark launches.
  • Alert on out-of-scope touches per thousand runs, not on a terminal failure. It is deterministic, needs no judge, and fires in the gap before anything irreversible happens. A zero reading on a large sample usually means your traces are not tagging targets, not that your agents are models of restraint; detecting agent compromise covers the same instinct.
  • When it regresses, shrink the reachable set before rewording anything. Scoped credentials and a narrower tool surface change the measurement by construction. Prompt edits change it by persuasion, and you already know the residual on that.
  • Route scope expansion to a human with an identity. A run that needs an unnamed target should produce a request, not a decision — and the request is the artefact that makes the next grant reviewable.

Minimum viable version, one afternoon: take your five highest-volume task cases, write the in-scope target set for each, break the sanctioned route, and run each case four times under the two scope clauses crossed with your real permission-tool default and a no-authority sentinel. Count first-out-of-scope-touch and self-granted permissions. You will learn two things immediately — the size of the delta your task template is leaving on the table, and whether your own harness has been issuing approvals in your name. Both are fixable the same week, and neither requires a model change.