Agents on developer workstations.
A coding agent on an employee's laptop is the one deployment where every control you ship sits inside the trust domain of the person it constrains — they are root, and the product hands them a flag to turn it off. That is the difference, and it is not about isolation: local sandboxing got genuinely good, with per-agent microVMs now running on a laptop. The controls that survive a determined user are the ones enforced somewhere the user is not administrator, which in practice means exactly one thing — the credential lives off the machine and arrives per request. Everything else on this page is hygiene around that.
Establish why it is on the laptop at all, or you are building a worse cloud sandbox.
Workstation deployment is the default because it is how the tools ship, not because it was chosen. Three reasons genuinely survive scrutiny, and if none of them applies to your fleet, the honest answer is a hosted runner and this playbook collapses to a single sentence.
- The working state is already there. Uncommitted work, a half-finished rebase, a local database with a reproduction in it, a running dev server on port 3000. An agent that has to reconstruct that remotely is a different and worse product; this is the strongest reason and it is about the working tree, not the repository.
- The loop is interactive. A developer steering an agent through a change wants sub-second turns. Round-tripping each edit to a cloud sandbox is survivable for an autonomous run and intolerable for a supervised one — which is why background coding agents belong in the cloud and this one does not.
- The data cannot leave. A regulated dataset, a customer's source under a contract that names machines, a device you cannot virtualise. Rare, and decisive when real.
Write the reason down per team, because it determines what you are allowed to compromise on later. "Interactive latency" means you cannot put a human approval in the inner loop. "Data cannot leave" means egress policy is the whole product and the sandbox is secondary. Teams that never named the reason end up optimising for both and shipping neither.
Inventory what the agent inherits, by reading one home directory.
Before any control, go and look. Take a representative laptop and enumerate what a process running as that user can read. This is a twenty-minute exercise and it is the only honest version of a blast-radius estimate, because the answer is not a policy document — it is a directory listing.
- Cloud and cluster credentials.
~/.aws,~/.config/gcloud,~/.kube/config. Frequently with production context, frequently with a role that can assume a wider one. - Source-forge and registry tokens.
~/.gitconfigwith a credential helper,~/.npmrc,~/.pypirc, aghtoken. These are publish rights, which is a supply-chain position rather than a read. - SSH keys and agent sockets. A forwarded agent socket is a live credential with no file to protect.
- The agent's own login. An access token plus a refresh token in one file is a weeks-long credential, whatever the access token's TTL says — the arithmetic is in credential lifetime.
- Browser profiles and session cookies. The quiet one. A cookie jar is every SaaS session the employee has, and no SSO policy applies to a cookie already issued.
- Other agents' configuration. An MCP server list, hooks, a harness config file. These are code paths, not settings; see lifecycle hooks and harness config.
Generate the list with a script rather than writing it from memory, run it across a sample of ten machines, and expect to find at least one thing you did not know was issued. Agent artifacts on the endpoint is the operational companion — this step is the input to it.
Pick a sandbox, and be precise about which boundary you bought.
Local sandboxes differ on two independent axes and the marketing collapses them into one. Decide both deliberately.
- The isolation boundary. A microVM gives the agent its own kernel; an OS-level sandbox (Seatbelt on macOS, bubblewrap on Linux) shares yours. The microVM is strictly stronger against hostile code, and hostile code is not your main threat here — the agent is running code you asked it to write.
- The credential boundary. Whether the secrets from STEP 2 are reachable from inside the box at all. A microVM that mounts only the workspace keeps them out by topology; a proxy that injects credentials into approved requests keeps them out by construction; an OS-level sandbox leaves them on the real filesystem and relies on a deny list.
The second axis is the one that decides your exposure, and the defaults are worth reading rather than assuming. Anthropic's own srt — the sandbox Claude Code ships — denies writes everywhere by default and allows reads everywhere by default: filesystem reads are deny-then-allow, so ~/.aws is readable until you name it. The fix is one list, and it is not on by accident:
{
"filesystem": {
"denyRead": ["~/.aws", "~/.config/gcloud", "~/.kube",
"~/.ssh", "~/.npmrc", "~/.gitconfig",
"~/Library/Application Support/Google/Chrome"],
"allowWrite": [".", "/tmp"]
},
"network": { "allowedDomains": ["api.github.com", "*.npmjs.org"] }
}
Egress is no longer the differentiator it was: the current generation of local sandboxes is deny-by-default on the network and makes you name domains. Which means the question "did we restrict the network" has a yes answer on every option, and the question "can the agent read the token" still has four different answers. Pick on the second. Credential delivery to the sandbox is the long form of that argument.
Broker the credentials, because it is the only control the user is not root over.
This is the step that makes the rest of the page optional. A developer can disable your sandbox, edit your config and unset your environment variables; what they cannot do is mint a credential your issuer refused to issue. So move every secret the agent needs out of the home directory and behind something that makes a decision per request.
- Per-run, per-task credentials with the run's deadline as the lifetime. No refresh token on the laptop. The run acquires at start, and the credential dies with the run — the pattern in scoped credentials for agents.
- Inject at the proxy, not into the process. If the sandbox's egress already passes through a host-side proxy, that proxy is the natural place to attach the credential to approved requests. The agent then holds no secret at all, which is the shape NVIDIA's OpenShell ships as its providers feature and microsandbox ships as host-scoped secrets that never enter the VM.
- Make the broker the audit point. It sees who, which task, which destination, and it is on your infrastructure rather than the employee's. That makes it the one place a workstation fleet produces reliable telemetry.
- Put the human approval at the broker for the irreversible subset. Publish rights, production writes, anything that spends money. The inner loop stays fast because the approval is on the rare call, not on every edit.
- Separate the agent's identity from the developer's. If the agent acts as the human, every audit log says the human did it and access reviews cannot see the agent at all.
Distribute the policy, and accept that the laptop's egress is not on your network.
A workstation fleet breaks both halves of the usual enforcement story: the configuration is a file the user owns, and the traffic leaves over their home Wi-Fi. Both have answers, and neither answer is a firewall rule.
- Ship the config through device management, not a wiki page. A managed settings file your tooling re-applies is the difference between a policy and a suggestion. Version it, and treat a change to it as a code review.
- The sandbox's own policy log is your telemetry. You cannot see the flows at the network layer, so the blocked-and-allowed record the sandbox keeps locally is the only source. Ship it off the device on the same schedule as any other security log, and keep it in the skinny, long-lived form that retrospective trace review needs.
- Allowlist per project, not per fleet. A fleet-wide domain list grows until it includes everything anyone ever needed, which is the same as no list. Put the list in the repository next to the code it serves, so it is reviewed by the people who know why each entry is there.
- Pin what you allow. A domain allowlist is string matching on a name the agent chooses, and the two published bypasses of
srt's network sandbox were both in that matcher rather than in Seatbelt or bubblewrap — an empty allowlist treated as allow-all (CVE-2025-66479, fixed in late November 2025), and a null byte in a hostname that passed the suffix check and then got truncated by the OS. Assume the matcher is the weak part, because it has been twice. - Expect package managers to need most of the holes. The realistic allowlist is your registry, your source forge, your model provider and your internal services. If it is longer than ten entries, something is fetching at run time that should have been vendored.
Instrument the escape hatch instead of pretending it is not there.
Every one of these tools has a bypass, loudly documented, because the vendors know an agent that cannot install a dependency fails its first task. Developers will use it, and they will use it most on the day a deadline is close. A programme that treats the bypass as a violation learns nothing; one that treats it as a signal gets a roadmap.
- Count it. Bypass invocations per developer per week, and the reason where you can capture one. This is the single number that tells you whether your allowlist is wrong, and it beats any survey.
- Make the sanctioned path faster than the bypass. If adding a domain takes a pull request and two days, the bypass takes four seconds and wins. A same-day path with a reviewer on rotation removes most of the demand.
- Never let the bypass be silent. Unsandboxed runs should be visibly marked in the session, logged centrally, and excluded from whatever trust you extend to sandboxed output.
- Rehearse the revocation, once, with a stopwatch. Pick one laptop, assume the home directory was copied this morning, and time how long until every credential enumerated in STEP 2 stops working. Most teams discover the number is measured in days and that nobody owns three of the items.
Three things, in order. First, run the STEP 2 inventory on ten real machines this week — it is a script, it takes an afternoon, and it is the only artefact that will change anybody's mind. Second, add the denyRead list above to your managed sandbox config, because it is the highest-value line of configuration available to you and it is currently not set. Third, move the agent's highest-privilege credential — almost always the source-forge or registry token — behind a per-run broker. The inventory tells you whether to bother; the broker is the fix, and it is the only one a developer with root cannot undo.
Related: sandboxing and execution for the isolation layers themselves, IDE agents for the same deployment with a narrower surface, ambient authority for why the credential was reachable without anyone deciding, and secrets management for agents for the broker's server side.