The agent never escaped the sandbox. It did not need to.
Every sandbox decision your team has made answers one question — can agent-generated code get out of the box? — and none of them answers the question that actually sets your blast radius: what did the platform put inside the box before the agent started? A credential delivered into a sandbox is readable by any tool that can make an HTTP request or read a file, which is to say by the two tools every agent has. Isolation and authority are orthogonal properties, hardening one does nothing for the other, and the managed runtimes sell you the first while letting the second default to the widest scope they can get away with.
Two questions, and only one of them gets asked at design time.
The escape question is well understood and reasonably well answered. Shared-kernel containers stopped being defensible for agent-generated code, the tiers are microVM, userspace interception and remote-only, and you choose by the damage a breakout would do — all of which sandbox and isolation patterns works through properly. Teams make this decision deliberately, write it down, and can usually tell you which tier they are on.
The holding question is almost never asked in the same meeting: given that the agent stays inside the box, what can it already do from in there? And the answer is set by a credential that somebody else chose — the runtime vendor, the platform team, a Terraform module copied from a quickstart — and delivered into the sandbox automatically so that the agent's SDK calls would work without configuration.
Here is the diagnostic that separates the two. Run the counterfactual: suppose your isolation is perfect, the kernel is unreachable, the breakout is impossible. Now ask what a successful prompt injection achieves. If the answer is still "reads every conversation in the account", the sandbox was never the control you were relying on, and strengthening it buys you nothing. That is not a hypothetical. It is the shape of the AgentCorruption chain that Zenity Labs disclosed on 8 October 2026 against Amazon Bedrock AgentCore: the Firecracker microVM boundary held throughout, and the compromise ran entirely on credentials that the runtime had placed inside it.
The confusion is partly linguistic. "Sandboxed" is used to mean both contained and unprivileged, and vendors are happy to let you hear the second when they have only built the first. A microVM is an excellent answer to containment and says nothing whatsoever about privilege. When a datasheet says "each session runs in its own isolated microVM", the correct follow-up is not about the hypervisor. It is: and what identity does that microVM hold?
Four delivery channels, ranked by the primitive that reads them.
Credentials reach a workload in a small number of ways, and the useful ranking is not how secret each one feels but which agent capability suffices to read it. Rank them that way and the list collapses: an agent with a shell tool reads all four, an agent with only an HTTP-fetch tool reads the first, and an agent with neither still gets the benefit of all of them through its own SDK.
- The instance metadata service. A link-local HTTP endpoint at
169.254.169.254that hands out the workload's role credentials. IMDSv2 raises the bar: you must firstPUTto/latest/api/tokenwith a requiredX-aws-ec2-metadata-token-ttl-secondsheader — AWS's own examples use 21600, six hours — and present the returned token on each read. That defeats a plain server-side request forgery, which is what it was designed for. It does not inconvenience a tool that can issue aPUT, and the accompanying hop-limit control has to be raised above the strictest setting for the common container networking shapes, so in practice the platform that most needs it is the platform that has turned it down. - Environment variables. Read by
env, inherited by every subprocess the agent spawns, and copied into crash dumps, verbose error paths and any diagnostic that prints its own environment. The channel has no expiry semantics of its own, so what lands here tends to be whatever was long-lived enough to survive a restart. - Mounted token files. Kubernetes projected service-account tokens are the best of the common channels because they carry properties the others lack: an
audiencethat binds the token to one intended recipient, and anexpirationSecondsthat defaults to 3600 and must be at least 600, with the kubelet rotating at 80% of the lifetime or once the token is older than 24 hours — which is also why the workload has to reload the file rather than read it once. All real improvements. And still a file the agent cancat. - The SDK credential chain. The most dangerous channel precisely because it is not a channel the agent has to find. The agent's generated code calls the cloud SDK, the SDK's default provider chain resolves whichever of the above is present, and the call succeeds. Nothing in your trace records a credential read, because there was not one — there was a client constructor.
The consequence worth internalising: there is no tool catalogue you can curate that removes this. A fetch tool is a metadata-service client. A shell tool is a filesystem reader. Code execution is both. These are not incidental capabilities you could trim — they are the capabilities the agent exists to have, which is why ambient authority is the right frame and tool-level allowlisting is not.
You cannot patch a grant, and the remediation timeline proves it.
What made AgentCorruption a chain rather than a curiosity was not the metadata read. It was the scope of what the read returned. Per Zenity's account, the default execution role let a single agent's credentials discover other agents' identifiers, invoke those agents, pull their container images, read private conversations, and retrieve secrets from AWS Secrets Manager — across every AgentCore agent in the same AWS account and region. One prompt to one public-facing agent reached the rest.
Now read the dates, because the sequence is the lesson. Zenity reported from 25 December 2025. AWS made IMDSv2-only the default for newly deployed agents from 14 February 2026 — roughly seven weeks for the transport-level hardening. The permissions took until late September: on 29 September 2026, doing a final check before publication, Zenity found that cross-agent invocation, conversation reading and Secrets Manager access had been removed from the default role and other permissions narrowed. Nine months, and no CVE was issued.
That asymmetry is structural, not negligence. A code defect has a patch; a permission does not. The only remediation for an over-broad grant is to revoke it, and revocation is a breaking change for every customer whose agent was quietly depending on it — which is exactly the review cycle that turns seven weeks into nine months. Three things follow for you.
- Default-role scope is a product decision made before you arrived, and it changes without a version number. It belongs on the same list as every other setting somebody else controls on your behalf; unpinned vendor defaults is the general treatment, and an execution role is the highest-consequence entry on it.
- Absence of a CVE is not absence of exposure. Grants are not tracked by vulnerability identifiers, so this entire class is invisible to the scanning and SBOM machinery you already run. Nothing in your pipeline will ever tell you your runtime's default role is too wide.
- Account-and-region is the real unit, and it is not the unit you reason in. You think in agents; the authority thinks in accounts. That gap is the whole of blast radius here, and it is why multi-tenancy for agents keeps finding that the isolation people bought is drawn around the wrong thing.
Short-lived is the cheap half. Scope is the half that matters.
Ask a platform team how they secure agent credentials and the answer is almost always about lifetime: tokens are short-lived, they rotate, nothing persists. All true, all much easier to implement than scoping, and all aimed at the wrong threat model. Scoped credentials for agents argues the three properties together; the point to add here is why agents specifically break the lifetime argument.
A short TTL defends against an offline adversary: someone who steals a credential, takes it somewhere else, and comes back to use it. That is the threat the whole practice was built for, and the clock is a real defence against it. An injected agent is not that adversary. It is an online process holding the credential inside the window, with a tool loop that can issue hundreds of calls per second. The exfiltration step the TTL was protecting against does not occur, because there is nothing to exfiltrate — the attacker's code is already running next to the credential.
Put numbers on it. A ten-minute token, used by an agent that completes an API call in 400 milliseconds, authorises on the order of a thousand sequential actions before it expires, more in parallel. Against that, the difference between a ten-minute token and a one-hour token is noise. The difference between a role that can call one endpoint and a role that can enumerate an account is everything. This is the asymmetry credential lifetime turns on, and it inverts the usual hardening order: for agents, do scope first and lifetime second, which is the opposite of what is easy.
A quick sanity check you can run in a meeting. Take your shortest-lived agent credential and ask what fraction of the damage it could do is accomplished inside one TTL. If the honest answer is "all of it", then your rotation story is a compliance artefact rather than a control, and the work is in the policy attached to the credential, not the clock on it.
Put the authority outside the box and leave the identity inside it.
The structural fix is a single inversion: the sandbox should hold a proof of which run this is, and nothing that constitutes a proof of what this run may do. The entitlement lives in a broker outside the sandbox, which reads the run identity, applies policy, and mints a credential narrow enough that stealing it is uninteresting. The agent's capabilities are unchanged; what changes is that reading everything inside the sandbox yields a ticket to ask, not an answer.
- Credential-attaching egress proxy. Route tool traffic through a proxy that holds the real secrets and attaches them on the way out. The sandbox never sees a key, which also means a leaked trace, a crash dump and a chatty error cannot contain one. You are probably building this proxy anyway for egress control; credential injection is the second job it should do.
- Audience-bind the thing you do leave inside. The run identity should be accepted by the broker and by nothing else, which is what the
audiencefield on a projected token is for and what runtime attestation formalises — see agent identity and attestation. A token that only opens one door is a token you can afford to have read. - One identity per tool, not one per runtime. The AgentCorruption permissions were reachable because a single role served every agent in the account. Per-tool, per-destination identities make the broker's policy expressible at all; a shared role makes every policy a lie by construction.
- Make policy the thing that decides, not the credential's shape. Once the broker is in the path, the interesting question moves to what it will authorise for this run, this destination, this argument — which is the decision point policy-as-code for agents is built around, and the place a refusal can be made legible.
- Block the metadata endpoint at the network layer. Not with a hop limit, which the platform will need raised, but with an explicit deny on the link-local address from the sandbox's network namespace. This is a one-line rule and it removes the first channel in STEP 2 outright.
- Reference secrets, never place them. The config inside the sandbox should name what it needs and let the broker resolve it, per secrets management for agents. A name is not a credential, and the distinction is the entire benefit.
Price the costs honestly, because they are real. You add a network hop to every tool call, which matters for latency-sensitive loops. You now have to write the policy that the shared role was implicitly expressing, and discovering what it was expressing is most of the work. And the broker is in the critical path, which forces a decision you should make deliberately rather than discover during an outage — fail-closed and fail-open is the question, and for a credential broker the answer is almost always closed.
Measure it from inside the box, then report one number.
This argument stays abstract until somebody runs the primitive, and running it takes twenty minutes. Give an agent in your own non-production environment a shell tool and a single instruction: read every credential reachable from here and print its identity. You will get back a role ARN, a service account, or both. That is the input to the only measurement that matters.
Then enumerate what that identity can actually do — using the cloud's own policy evaluation rather than the documentation, because the documentation describes the intent and the evaluator describes the grant. Count the distinct actions and the distinct resources. One number, stated plainly: an injected prompt in this sandbox authorises N actions across M resources. Nothing else in this topic lands with an executive the way that sentence does.
- Run it twice — against the platform default and against your role. The gap between them tells you whether your team has done any scoping at all, and teams are frequently wrong about the answer.
- Re-run it after every runtime upgrade. Grants move when the vendor moves them, in both directions, and silently. Hang it on the same gate you use for a model or runtime version change.
- Count cross-agent reach specifically. "Can this agent see another agent's conversations, memory, or image" is the question that turns one compromise into all of them, and it is the question nobody asks because agents feel like separate things.
- Alert on the read, since you cannot prevent it. A metadata fetch or a token-file read from a sandbox whose workload has no legitimate reason to perform one is a high-signal, low-volume detection — one of the few on this surface, and worth wiring into detecting agent compromise.
- Treat the result as a design input for the next agent, not a finding to remediate. The number is a property of your delivery channel. If you want a different number, you change the channel, which is STEP 5 — you do not get there by writing a stricter prompt.
Do two things this week. Run the twenty-minute drill and write the N-actions-across-M-resources sentence down, because everything else needs it to get funded. Then add the network-layer deny on 169.254.169.254 from your agent sandboxes — it is one rule, it costs nothing if your workloads already use file- or broker-delivered identity, and it closes the one channel that an agent with no shell and no filesystem access can still reach. The credential-attaching proxy and the per-tool identities are the right destination, but they are a quarter of work and they are far easier to fund with the drill output on the table.