The headline from OWASP's 2026 release is that Excessive Agency climbed from sixth place to third, and that is the least interesting thing in it. The change worth your afternoon is that the GenAI Security Project stopped shipping prose and started shipping interfaces: the Agent Control Standard defines the hook a framework fires before your agent calls a tool, writes to memory, or spawns a sub-agent, and the verdict — allow, deny, modify — your policy returns. Advice you could ignore has become a contract a framework either implements or does not, which moves the audit boundary from your policy binder into your runtime. It also creates a new number nobody is reporting: the fraction of your agent's actions that pass through a hooked call site at all.
At a glance
Two artefacts landed under the same banner, and they do completely different jobs.
| Artefact | Shipped | What it is | What it asks of you |
|---|---|---|---|
| Top 10 for LLM Applications 2026 | August 2026 | A ranked risk list, for the first time weighted against incident data | Read it; re-prioritise |
| Agent Control Standard (ACS) | Launched May 2026; under OWASP from September | Three interface layers: runtime hooks, trace conventions, an agent BOM | Implement it, or pick a framework that has |
| Agent Observability Standard (AOS) | Alongside, as a separate project | Instrumentation conventions for agent telemetry | Emit it |
The 2026 list, in order: Prompt Injection, Sensitive Information Disclosure, Excessive Agency, Supply Chain, Data and Model Poisoning, Unbounded Consumption, Misinformation, Hidden Context Exposure, Vector and Embedding Weaknesses, Improper Output Handling. The top two are unchanged, which is itself a finding after three years of the field insisting a fix was imminent.
What the list actually says
The methodology changed, and that is why the ranking moved
This edition is the first to weigh community judgement against a body of real incidents: the expert vote carries 75% and the remaining quarter comes from roughly 6,639 incidents drawn from public vulnerability databases and an AI-harm database. Excessive Agency is the entry where the two sources agree most strongly, and a three-place climb is what agreement looks like when one of the inputs is an incident count rather than a survey.
Read that as a statement about where failures are actually occurring. Excessive Agency is not a model weakness — it is the class where a system was handed more autonomy, permission, or unsupervised reach than the task needed, so that an ordinary mistake or a successful manipulation produces a disproportionate effect. That is a description of a deployment decision, and it has now been ranked above supply-chain and poisoning risks by the incident record. Our blast radius and ambient authority pages are the same argument from the other end.
The quiet rename is the one to act on
LLM08, previously system-prompt leakage, is now Hidden Context Exposure, and the scope broadened from the system prompt to all hidden operational context — retrieved documents, MCP tool schemas, formatting rules, whatever else your harness assembles and never shows the user. That is a much larger surface than the old entry, and most teams have controls for exactly one of its members.
It also lands on a real asymmetry. A leaked system prompt is embarrassing; a leaked tool schema is a map of your agent's capabilities, argument shapes and internal service names, handed to whoever asked for it politely. See system prompt extraction for what extraction looks like in practice.
What did not move
Prompt Injection is still first. Three years in, with every major vendor shipping a classifier and a guard model, the risk that heads this list is the one with no complete fix — which is the correct outcome, not a failure of the list. The control taxonomy that follows from accepting that is in prompt injection defence in 2026.
The part that is not a list
ACS is organised in three layers, and the separation is the useful part. Instrument defines the runtime hooks and the Guardian Agent pattern: a callback the agent's host invokes before an action executes, carrying enough context — the tool, the arguments, the identity, the prior turns — for a separate evaluator to permit, deny, modify, ask about, or defer it. Trace extends OpenTelemetry and OCSF with agent-specific semantic conventions, which is the same ground as the GenAI semantic conventions. Inspect extends CycloneDX, SPDX and SWID to emit a dynamic agent bill of materials — the thing every inventory obligation has been asking for and nobody could produce.
The design choice that makes this more than another framework is that the spec is policy-engine agnostic. It defines the hook contract and says nothing about what runs behind it; a framework-agnostic SDK reads declarative policy and enforces it through whatever hooks the platform exposes. If you have already invested in a policy engine, that investment survives — which is the argument in policy as code for agents, now with a defined place to attach.
Note also what this does to procurement. "Does your agent framework implement ACS Instrument?" is a question with a yes-or-no answer, unlike "how do you handle tool authorisation?", which every vendor can answer well. A hook contract is a capability you can test in an afternoon, and the frameworks that have no interception point at all cannot hide that behind a security page.
A hook is not a control
The failure mode available here is old and well documented under other names. A hook that fires, logs, and lets the action proceed is telemetry; it produces a dashboard of everything you did not stop. An advisory verdict that is computed but not enforced is worse, because it pays the latency and buys the appearance of the control. The property that makes a hook a control is that it is inline, deterministic, and blocking — the action does not reach the production system until the verdict returns.
Two hooks carry most of the weight, and they are not the two people instrument first. The pre-execution tool-call hook is the obvious one, and it is where authorisation belongs: identity, scope, argument bounds. The hook on the tool result is the one that matters for the risk sitting at number one, because injected instructions arrive in retrieved documents, web pages, ticket bodies and MCP responses — inbound, after the call you authorised. A guardrail that inspects requests and waves results through is inspecting the half of the channel the attacker is not using. That is the point of treating telemetry as untrusted input, generalised.
And the verdict vocabulary deserves more scepticism than it is getting. Allow and deny are easy to reason about. Modify means a policy layer is rewriting arguments or content before the agent sees them, which is a capability worth having and a debugging surface worth fearing: a run that behaved oddly because a guardian silently rewrote a parameter is nearly impossible to reconstruct unless the hook's own decisions are in the trace with the same fidelity as the tool calls. Log the verdict, the rule that produced it, and the before-and-after — or do not use modify.
The number nobody is reporting
A standard hook surface standardises the bypass as well. Every hook in the Instrument layer fires inside the agent's step, which means it covers actions the framework mediates and is silent on everything that leaves by another door:
- Code the agent wrote and then executed. A generated script that opens a socket or shells out makes its calls inside the sandbox, not through the tool interface. The
execute codehook fires once, on the execution; it does not fire per syscall. Code as action is a deliberate architecture, and it trades tool-level interception for expressiveness. - Network calls inside a tool implementation. Your hook sees
search(query). It does not see the four HTTP requests the implementation makes, one of which goes somewhere you would not have approved. - Anything outside the framework. A cron job with an API key, a notebook, a colleague's script — none of it has a host to fire a hook.
- Sub-agents in another runtime. The invoke-sub-agent hook fires at your boundary. What the sub-agent does inside a different framework is governed by that framework's hooks, if it has any.
So the metric that tells you whether any of this is working is a denominator, not a count: what fraction of your agent's externally visible effects passed through a hooked call site? A team that has implemented ACS, reports 40,000 evaluated tool calls a month, and has an agent that writes and runs Python has not measured the thing they think they measured. Start from the threat model and enumerate the exits before you celebrate the coverage.
What to do with each layer
| If your problem is… | Instrument (hooks) | Trace | Inspect (BOM) |
|---|---|---|---|
| An agent can take an action it should not | This is the layer | Tells you afterwards | No |
| Injected instructions in tool results | Result hook, enforcing | Evidence for triage | No |
| "Which agents do we run?" | No | Partially, from traffic | This is the layer |
| An auditor wants proof of control | The verdict log | The trajectory | The component list |
| Vendor selection | Ask for conformance | Ask for conventions | Ask for the BOM |
Implement Instrument on the two tool hooks first, enforcing rather than advisory, and put the verdicts in the trace. That is a week of work and it is most of the value. Inspect is worth adopting early for a different reason: it is the first credible answer to an inventory obligation that does not depend on people filing forms.
FAQ
Does ACS replace my guardrail or policy product?
No — it is the socket those products plug into. The specification defines the hook contract and deliberately says nothing about what evaluates policy behind it, so an existing engine keeps its rules and gains a standard attachment point across frameworks.
Is the Top 10 ranking change worth re-prioritising on?
The movement matters less than the reason for it. Excessive Agency rose because incident data agreed with the expert vote, which is evidence that real failures are coming from over-broad permissions rather than from novel model attacks. If your controls are concentrated on the model and thin on what the agent is allowed to do, that is the signal.
Why is the tool-result hook more important than the tool-call hook?
Because the request is the part you authored and the result is the part an attacker can write. Prompt injection arrives inbound, in retrieved text, page content, ticket bodies and MCP responses. Inspecting only outbound calls leaves the channel that carries the top-ranked risk uninspected.
Does implementing hooks satisfy an audit?
Only if the verdicts are logged and the coverage is stated. A hook that fires and lets everything through produces telemetry, not evidence of control, and an auditor who asks what fraction of actions were evaluated will not be satisfied by a count of evaluations with no denominator.
What breaks first when you turn enforcement on?
Latency budgets and false positives, in that order. Every mediated call now waits on a policy decision, and the first week of blocking verdicts will include legitimate actions. Run the policy in advisory mode long enough to size the false-positive rate, then switch to blocking — but do switch, because advisory mode has the cost of the control and none of the benefit.
Further reading
On this wiki:
- Policy as Code for Agents — what runs behind the hook.
- Blast Radius — Excessive Agency stated as a design parameter.
- Prompt Injection Defence in 2026 — the control taxonomy for the entry that did not move.
- OTel GenAI Semantic Conventions — the ground the Trace layer extends.
- Agent Inventory & Registry — what a dynamic agent BOM is finally an answer to.