The confused deputy.
Every control you installed to stop your agent doing something forbidden is useless against the attack where it never does anything forbidden. In a confused-deputy attack the agent is authorised, the action is permitted, the audit log is clean, and the damage is complete — because your agent supplied the authority while a web page supplied the target. The question that finds these is not "what is the agent allowed to do" but "where did the name of the thing it acted on come from", and if the answer is anywhere other than the same place the permission came from, you have one.
The 1970s bug that named the problem.
Norm Hardy's three-page note The Confused Deputy (or why capabilities might have been invented), published in ACM SIGOPS Operating Systems Review in 1988, describes a failure he had watched at Tymshare on a PDP-10. The Fortran compiler, (SYSX)FORT, collected statistics about which language features customers used and wrote them to (SYSX)STAT. To let it do that, administrators granted the compiler a home files license: permission to write inside the SYSX directory.
The compiler also accepted an optional filename from the user, for debug output. A customer who had noticed that the system billing log lived at (SYSX)BILL passed that name as the debug destination. The compiler opened it — using its own license, because that is the license it had — and overwrote the billing records with compiler diagnostics.
Read the structure rather than the anecdote. Every action has two inputs:
- Authority — the answer to "is this principal permitted to do this?" Here it came from the compiler's license.
- Designation — the answer to "which object?" Here it came from the caller's argument.
Nobody exceeded their permissions. The customer had no right to write the billing file and never acquired one; they simply arranged for something that did have the right to write it on their behalf. The deputy was not compromised. It was confused about whose errand it was running.
This is why the paper's subtitle is about capabilities. In a capability system you cannot name an object you have no right to — the name is the permission, so authority and designation arrive together and cannot be mismatched. Every mitigation in STEP 4 is a way of recovering that property in a system that was not built with it.
An agent is a confused deputy by construction, not by accident.
Hardy's compiler had to be handed a bad filename deliberately. An agent goes looking for them. Separate the two inputs again and the default architecture is alarming:
- Authority comes from you. The logged-in browser profile, the OAuth tokens in the MCP config, the cloud role on the container, the repository write scope. All of it is standing and unenumerated — this is ambient authority, and it is the reason the blast radius of a confused agent is "everything the session reaches" rather than "the twelve tools you registered".
- Designation comes from the content it reads. The URL in the issue comment, the path in the README, the account number in the PDF, the instruction in the HTML comment, the review text on the product page. The agent's whole value proposition is that it takes targets from material it was not able to vet.
Prompt injection is how designation gets smuggled in; confused deputy is why that matters. Keeping the two ideas distinct is practically useful, because it tells you that reducing injection success rates and fixing the vulnerability class are different projects — and that the class predates language models by three decades. You already know its other names:
- CSRF. The browser is the deputy, your session cookie is the authority, and the attacker's page supplies the designation.
- SSRF. The server is the deputy, its position inside the network is the authority, and a user-supplied URL is the designation — which is exactly the shape of an agent with a
fetchtool. - Token passthrough. A gateway that forwards the token it was given to an upstream it was not issued for is a deputy presenting someone else's authority against a target it chose.
The uncomfortable corollary: a model that is working perfectly is the most dangerous deputy, because the attack needs it to be competent and obedient, not broken. See the principal-agent problem for the version of this that is about incentives rather than naming.
Four controls that feel like the fix and are not.
Each of these is worth having for other reasons. None of them addresses a deputy acting within its permissions.
- Stronger agent identity. Workload identity, short-lived certificates, signed requests — all good, all orthogonal. The deputy was already authenticated. Proving more rigorously which deputy it is does not tell the resource whose errand it is on.
- Confirmation dialogs. A prompt that says "allow the agent to write to
invoices/bill.csv?" is asking the user to ratify the attacker's designation. The dialog can only present the action the agent asked for, and the agent asked for the wrong thing sincerely. Worse, approval fatigue makes the hundredth dialog cheaper than the first. - Prompt hardening. System-prompt rules and the instruction hierarchy lower the rate. They cannot change the class, because the model is being asked to tell data from instruction in a channel where the two are the same bytes. A rate reduction is risk management, not mediation.
- Audit logging. The log will faithfully record an authorised principal performing a permitted action on a resource it has rights to. It is a true record of a successful attack, and the reference-monitor test says why: a control downstream of the confusion inherits it.
A useful tell: if your mitigation would still be satisfied when the attacker picks the target, it is not a mitigation for this. Run that sentence against each control above.
The fix is to make authority travel with designation.
You cannot give the model better judgement about which filename to trust. You can make the untrusted filename not be a grant of anything. Four shapes do that, in rising order of effort:
- Hand out handles, not strings. A tool whose parameter is an opaque identifier minted by your own code for one object cannot be pointed at a second object by anything the agent read. A tool whose parameter is a path, a URL or an account number can. This is the single highest-leverage change in tool design, and it is a schema decision, not a security feature.
- Scope the credential to the task, not the agent. One pre-signed URL for the one file, valid for minutes; a token whose audience is the single upstream it will be presented to; a repository token with write access to one branch. The agent that holds nothing broad has nothing broad to be talked into using.
- Exchange tokens, never forward them. At every hop, mint a new credential bound to the real end user and the real destination. A forwarded token is the billing-file license, repeated at every layer.
- Split read authority from write authority. Two identities, two network paths. The identity that fetched the untrusted page should not be the identity that can act on what the page said. This is the architectural version of the same move, and it is what makes the blast radius of a successful injection finite.
Where you cannot do any of that — a legacy tool that takes a path and a credential that cannot be narrowed — fall back to constraining the designation itself: an allowlist of targets computed before the untrusted content is read, enforced outside the model. That is weaker, but it is still mediation, and it is strictly better than asking the agent to be suspicious.
Do this audit this afternoon, because it takes an hour and it is the one that finds real bugs. Make a table of every tool your agent can call. Two columns: where does the authority come from, and where does the target come from. Any row where those are different sources is a confused deputy, and the rows where the target comes from fetched content are the ones to fix first. Most teams find their worst row is a fetch, a file write, or a database query that takes a free-text identifier.