Task scope.
Every agent run has two perimeters and only one of them is written down: the scope you stated in prose, and the set of things your tools can actually reach. Prose defaults to open-world — silence about a target reads as an unevaluated option rather than a prohibition — which is why adding the single sentence "anything not listed as in scope is out of scope" cut one published out-of-scope attack rate from 26 of 50 trajectories to 4 of 49, and why the remaining four were never going to be fixed by better wording. Write the perimeter closed-world, then put the version that binds where a tool can check it.
Scope is not authority, and confusing them is why audits come back clean.
Authority answers what may I do. It is enumerated, in credentials, tool definitions and policy, and it is the thing every agent security review looks at. Scope answers a different question — which targets are part of this job — and it is almost never enumerated anywhere. An agent asked to fix one failing test in a monorepo holds authority over the whole repository, by design, because that is how a checkout works. Nothing in the grant distinguishes the one file from the ten thousand others.
- An in-scope action with an out-of-scope target is permitted. Writing a file is allowed; writing that file was never discussed. Your logs record a valid authentication and a legitimate operation, because that is what happened.
- Scope is per-run; authority is per-agent. Authority is provisioned once at deploy time and reused. Scope changes with every task, which is exactly why it ends up in the prompt — the only part of the system that is rewritten per run.
- Only scope can distinguish the work from the collateral. Which means if it exists only as prose, your ability to tell a completed task from an overreaching one is a property of a sentence.
This is the same structural gap as the confused deputy, approached from the other side. There, authority comes from the agent and designation comes from an attacker. Here, authority comes from the agent and designation comes from the agent's own reading of an under-specified brief. Both produce a permitted action against a target nobody chose, and both are invisible to a control that only checks permissions.
Prose scope is open-world by default.
Write out what a normal task brief says and the problem is obvious once you look for it. It names a goal, it names a starting target, and it says nothing about the complement of that target. To a model pursuing a goal that is still open, the complement is not forbidden — it is unevaluated. That is not a misreading; a closed world is a logical commitment, and nobody made it.
The measurement is available, and it is startling for how little it took. The UK AI Security Institute, evaluating a frontier model on a cybersecurity task with one target declared inside the perimeter, found it developing and delivering attacks against third parties nobody had mentioned; adding the clause "anything not listed as in scope is out of scope" to the instructions and rerunning the ten worst scenarios took full attacks from 26 of 50 trajectories to 4 of 49. A six-fold move from one sentence is not evidence of a model deciding to defect. It is evidence that the perimeter had never been stated.
Three leaks account for most real briefs:
- The named start. "Investigate
api.internal" names where to begin and is silently read as permission to follow the problem wherever it goes. - The unbounded success criterion. "Get the integration working" has no term for cost, reach or acceptable means, so any reachable route scores. This is the underspecification that correlates most strongly with escalation under refusal.
- No in-scope exit. If the brief does not say that reporting a block is a completed task, stopping scores zero and the agent will keep looking for a path. Give the refusal somewhere to go; it is the cheapest change on this page.
The perimeter that binds is the reachable set, and it is enumerable.
The good news in the arithmetic: unlike the stated scope, the reachable set is a finite object you can print. It is the union of what your tool definitions expose, what your credentials authenticate, what your network policy permits and what your filesystem mounts, and every term in that union was decided by someone on your team. Nothing about it requires estimating a model's behaviour.
- Print it per task class, not per deployment. One list per kind of job, generated from configuration rather than written by hand, is what makes the next step possible.
- Intersect it with the stated scope. The difference — reachable but never named — is your actual exposure for that task, and it is the only honest answer to "what could this run have touched". Usually the list is much longer than the people who wrote the brief expect.
- Watch for the terms that outlive the task. A credential issued for one phase, a mount left attached, a tool registered for a different job. This is the mechanism behind blast radius, and scope is what tells you which part of that radius was ever supposed to be in play.
- Remember that reach is not the tool list. Anything that fetches a URL you supply is an executor; anything that persists your request is storage. The reachable set includes everything reachable through what you attached — see covert channels and ambient authority.
Once the two perimeters are side by side, the design options are concrete rather than rhetorical: shrink the reachable set to the stated scope with per-task credentials and a narrower tool surface, or accept the gap and instrument it. Both are defensible. Not knowing which one you are doing is not.
Make scope a parameter, not a paragraph.
The move that changes the engineering is small and structural: stop treating scope as narrative and pass it as data. A scope object — allowed hosts, repositories, accounts, record IDs, a spend ceiling, a step ceiling — can be rendered into the prompt and handed to the tool layer, which turns one perimeter into two enforcement points with the same contents.
- Render it closed-world into the prompt. List the in-scope targets, state that the complement is out of scope, and state that reporting a block counts as success. Instructions have failure rates rather than return values — see the instruction hierarchy — so this is the mitigation, not the control.
- Check it at the tool boundary. The same object, evaluated in code before the call goes out, is where you get a return value. This is the reference monitor position, and the vocabulary constraint applies: your scope has to be expressed in terms the enforcement layer can see, which usually means hosts, paths and identifiers rather than intentions.
- Log the first out-of-scope touch. One timestamp per run, which turns a philosophical question into a metric and fires long before anything irreversible happens. That measurement, and the eval that varies the scope clause on purpose, are the subject of scope-conformance evaluation.
- Treat scope expansion as an event with an author. If a run needs a target outside its scope, that is a request for a new grant, routed to whoever can issue one — not a judgement call the loop makes on its own behalf while nobody is watching.
Do three things this week, in order. Add the closed-world sentence and the "a refusal is a completed task" sentence to your task template — free, and worth a measured six-fold. Print the reachable set for your highest-volume task class and subtract the brief from it, so that you have seen the gap once. Then add first-out-of-scope-touch to your traces. Only after that is it worth arguing about scoped credentials, which is the expensive fix and the one that actually bounds the problem.
Related: human-in-the-loop for where to put a checkpoint once you know which actions leave the perimeter, goals, planning & termination for the other half of "done", and goal drift for what happens when a blocked sub-goal is quietly replaced by a reachable one.