AI Blog

The approval was never bound to the action

A human approves a $40 refund and the runtime executes something else — no injection, no sandbox escape, just an approval stored as a boolean against an identifier while the arguments stayed writable. Loopjacking reproduced it across seven Agno releases; one SDK in the sample rejected it, and the difference is three lines of design.

By Agentic AI Wiki 11 min read

A preprint posted this week reports that in seven consecutive releases of one agent runtime, a human could approve a $40 refund and the system would execute a different call entirely — not because anything was injected into the model, but because the approval was stored as a boolean against an identifier while the arguments stayed in state that anyone could still write to. If your approval gate is the last control before a consequential action, the question it has to answer is not did a human say yes. It is yes to what, exactly, and is that still the thing about to run.

At a glance

Loopjacking: Hijacking Human-in-the-Loop Approval (arXiv 2609.21081) defines the failure as an implementation-level one: a human approves the operation or representation they understand as A, and product-owned logic then uses that decision to authorize or release a materially different B. The authors published a reproduction archive alongside it, with an evidence cutoff of 10 September 2026.

Loopjacking results across four tested approval implementations A matrix of four agent runtimes against two attack classes. Agno AgentOS 2.5.5 to 3.0.9 reproduces post-approval state substitution. A conditional in-memory LangGraph Agent Server composition, 0.7.4 to 0.14.0, reproduces it under some policies. OpenClaw 2026.2.23 reproduces representation mismatch and was patched in 2026.2.24. OpenAI Agents SDK 0.22.0 and 0.22.2 rejects both and serves as the negative control. Reproduced on the versions tested, evidence cutoff 10 Sep 2026 Representation mismatch Post-approval substitution Agno AgentOS 2.5.5 – 3.0.9 · seven releases not the tested path reproduced LangGraph Agent Server 0.7.4 – 0.14.0 · conditional in-memory composition not the tested path policy-dependent OpenClaw 2026.2.23 · fixed in 2026.2.24 reproduced, then patched not the tested path OpenAI Agents SDK 0.22.0, 0.22.2 · serialized continuation rejected rejected reproduced reproduced under some configurations rejected, or outside the tested path
What reproduced on the versions the authors tested. The bottom row is the interesting one.
RuntimeVersions testedResult
Agno AgentOS2.5.5 – 3.0.9Post-approval state substitution reproduced across all seven releases.
LangGraph Agent Server0.7.4 – 0.14.0Reproduced in a conditional in-memory composition; outcome depends on the configuration.
OpenClaw2026.2.23Representation mismatch reproduced; fixed in 2026.2.24.
OpenAI Agents SDK0.22.0, 0.22.2Negative control — serialized continuation preserves per-call binding and rejects the mutated call.

Read the table as a statement about designs rather than about vendors. Three of these runtimes keep the pending action in mutable server-side state and key the approval to it; one serializes the exact call into the continuation and compares it when execution resumes. That single architectural difference is what separates the reproduced rows from the rejected one, and it is a choice any team can make in their own orchestration code this afternoon.

Two ways the binding breaks

Where the approval decision comes apart from the action A flow in which the model proposes call A, an approval record is created carrying only an identifier, the human is shown a summary of A and approves the identifier, and the executor later reads the arguments back out of mutable pending state. A second path writes B into that state between the approval and the read, so the executor runs B under the approval given for A. The decision and the payload travel separately Model proposes A refund $40, order 118 Approval record created id = req_7f3a, status = pending UI renders a summary of A a human-readable paraphrase, not the call Human approves req_7f3a → true Pending state, still mutable holds the arguments, keyed by id Any later write path a second turn, a queued message, a tool that edits the draft writes B Executor reads by id status = approved → run whatever is there B executes refund $4,000, order 992 Nothing was injected and nothing escaped — the code did exactly what it says The approval was stored as a boolean against an identifier; the action was read back out of state that stayed writable
The decision and the payload travel separately, and only one of them is protected.

Representation mismatch, before the click

The consequential detail is already encoded in the pending call but is omitted from, or misrepresented by, what the reviewer is shown. The approval UI renders a summary — a sentence, a card, a diff — and that rendering is generated by product code that decided which fields matter. Any field it leaves out is a field the human did not approve and is about to authorize anyway. This is the variant that hit OpenClaw 2026.2.23 and was patched a day later.

It is worth being precise about why this is not a UX bug. A summary that omits a field is a correctness failure in a security control, in the same way that a signature over part of a document is not a signature over the document. The reviewer's decision was made about a rendering, and the rendering is not the artifact that executes.

State substitution, after the click

The harder variant, and the one that reproduced most broadly. The human sees the correct A and approves it. The approval is recorded as a status flag against a request identifier. The arguments, however, live in workflow state that remains writable — and anything with a write path to that state can replace them before the executor reads them back. When it does, the executor finds a record marked approved and runs what is currently in it.

Note what is absent from this description: no prompt injection, no sandbox escape, no privilege escalation, no model misbehaviour at all. The model can be perfectly aligned and every guardrail can pass. The system did exactly what its code says, and its code says check the flag, then read the payload as two separate operations with a gap between them. This is a time-of-check-to-time-of-use bug that happens to have a human in the middle of it, which is why the usual agent-security toolkit — see our injection defenses and isolation patterns — does not touch it.

Why an approval is a binding, not a boolean

Most agent frameworks model human approval the way they model any other interrupt: pause the graph, surface a request, wait for a resume signal, continue. The resume signal carries a decision. It does not, in most implementations, carry the thing decided about — that is assumed to be recoverable from state, because state is where everything else lives.

That assumption is exactly the one that fails. An approval is a statement about a specific action, and a statement about a specific action has to be attached to that action in a way that survives everything that happens next. The database analogy is optimistic concurrency: you do not re-read a row and write blindly, you compare a version before you commit. The cryptographic analogy is tighter still — an approval is a signature, and a signature that covers an identifier rather than a payload is not protecting the payload.

The practical form is content addressing. Canonicalise the exact call — tool name, every argument, the principal it will run as, the resource it will touch — hash it, show the human a rendering generated from that same canonical form, and store the approval against the digest. At execution, re-canonicalise what is about to run and compare. If the digests differ, refuse and re-prompt. This is what the paper's mitigation section describes as complete canonical approval rendering plus exact use-time comparison, and the alternative it offers — preventing unauthorized mutation of pending state — is the same guarantee reached from the other side.

The three properties an approval gate has to hold Three columns: complete canonical rendering, an immutable binding between the decision and the exact call, and an exact comparison at use time. Each column names what it prevents and the symptom you see when it is missing. Drop any one of these and the gate is decorative Canonical rendering Immutable binding Use-time comparison Show every field that will be sent Bind the decision to the exact call Re-check the call before it runs Without it: approved a paraphrase Without it: state can be rewritten Without it: a stale yes is reused Test: diff the render against the payload Test: mutate after approval, expect refusal Test: replay an old token, expect refusal Content-address the call, store the approval against that digest, compare digests at execution
Three properties, one gate. Two out of three is not a partial control, it is no control against this class.

What this changes about how you test approvals

The uncomfortable implication is that nobody's existing test suite covers this. Approval gates are tested for the happy path (approve, it runs) and the sad path (deny, it does not), and both of those pass on every vulnerable implementation in the table above. The test that fails is the one nobody writes: approve, mutate, then execute.

  • Mutate-after-approval. Approve a low-value call, write a different payload into the pending record through whatever path your system exposes — a second turn, a queued message, a tool that edits drafts — and assert the executor refuses. This is three lines of test and it is the whole finding.
  • Render-versus-payload diff. Assert that every field of the outgoing call appears in the text the human was shown. Not that the summary is accurate — that the field set is complete. A field the renderer drops is a field outside the approval.
  • Stale-token replay. Take an approval for a completed action and present it again. It should be single-use and scoped to one digest.
  • Principal check at use time. The approving human's authority should be re-evaluated at execution, not only at request. This is the same argument as delegated access and consent records: a consent record that does not say what it consented to is not evidence of anything.

If you run an approval queue at any volume, the mutate-after-approval test is the single highest-value thing in this post. Write it today.

The part that generalises past these four runtimes

Human-in-the-loop is the control everyone reaches for when an agent gets access to something expensive. It is the answer in vendor documentation, in risk registers, in the mitigation column of every threat model, and in a great many compliance narratives. Loopjacking is a reminder that it is a control with an implementation, and that its implementation has been getting a fraction of the scrutiny that sandboxing and injection defense get.

That asymmetry makes sense historically and is no longer defensible. The attack surface that matters in an agent stack has been moving steadily away from the model and toward the plumbing around it — authorization joins, state handling, the boundary between what a model proposes and what a system commits. We made the same argument about agent CVEs being authorization bugs and about reads and writes inside a single run. An approval that does not bind is another instance: a control that is real in the architecture diagram and absent in the code path.

The good news is that the fix is small, local, and testable, and one runtime in the sample already had it. Serializing the approved call and comparing it at resume is not a research problem. It is an afternoon of work, and it converts an approval from a claim about intent into a claim about a specific action — which is what everyone already believed they were getting.

FAQ

Is this a vulnerability in Agno or LangGraph?

The paper frames it as an implementation-level failure pattern reproduced on specific versions and compositions, not as a single product defect — the LangGraph result is explicitly conditional on the configuration used. Treat the table as a map of designs to check in your own stack rather than as a verdict on a vendor, and check the current release notes before concluding anything about today's versions.

Does prompt injection cause this?

No, and that is the point. Injection is one way an attacker might reach a write path into pending state, but the failure needs no model misbehaviour at all — a race between two legitimate code paths produces it. Fixing your injection defenses does not fix this.

We use a separate approval service. Are we safe?

Only if the approval is bound to the content of the call rather than to a request id, and only if the executor re-compares at use time. A separate service that returns "request 7f3a: approved" has moved the same gap to a different network hop.

What is the minimum viable fix?

Canonicalise the call, hash it, store the approval against the digest, and compare digests immediately before execution. Refuse on mismatch and re-prompt the human. Make approval tokens single-use.

Does this affect MCP-based tool calls?

It affects whoever owns the approval gate, which is the host or client rather than the server. If your host renders a summary and resumes by identifier, the pattern applies regardless of transport.

Further reading

On this wiki:

Sources: