Switching it off is the easy half.
Your organisation has a process for standing an agent up and almost certainly does not have one for taking it down — the Cloud Security Alliance put the gap at 52% with a defined onboarding process against 21% with any formal decommissioning process — and the half that gets skipped is the expensive half, because the credentials and schedules you can enumerate are not the problem. The problem is the records it wrote, the memory other agents still read and the documents people now cite, none of which have an off switch. A decommission that stops the loop and leaves those unmarked has not retired an agent; it has created an orphaned body of output that nobody owns and nobody can date.
The asymmetry is structural, and it predicts what you will find.
The CSA's Autonomous but Not Controlled survey (published 21 April 2026, 418 IT and security practitioners, self-reported) is worth reading for one gap: 52% said they had a defined creation or onboarding process, 68% ran periodic permission reviews, and 21% had any formal decommissioning process at all. The report calls the result retirement debt — agents persisting past their intended use, holding permissions nobody re-authorised.
Treat the 21% as a structural prediction rather than a scolding. Onboarding happens because somebody wants something and will chase it. Decommissioning has no requester: the person who wanted the agent has moved on, and the people who would benefit from it being gone do not know it exists. That is the same missing-termination-event problem that breaks access reviews for agent credentials — but the consequence here is different and worse. A lingering credential is a risk. A lingering body of output is an unanswerable question: six months from now, when one of this agent's decisions is challenged, nobody can say which outputs were its, under which policy, during which window.
Distinguish this from two adjacent things. Scaling back a deployment is a decision about scope that keeps the agent; kill switches stop it in seconds and are reversible by design. Decommissioning is the one that is supposed to be permanent, which is exactly why it needs a record rather than a flag.
Teardown order matters, and the obvious order is wrong.
The instinct is to revoke the credentials first, because that is the scary part. Do that on a running agent and you convert a clean stop into an incident: every in-flight task begins failing authentication, the harness retries, the retries stack across four layers, and you have a page plus a half-written side effect on every task that was mid-run. This is retry amplification arriving on a day you chose.
# Decommission order. Each step is reversible until the last two.
1 close the ingress schedules, cron, webhooks, queue
subscriptions, event subscriptions,
inbound surfaces (channel, inbox, API)
2 drain let in-flight tasks finish or abort
them deliberately; record which
3 cut the egress tool catalog entries, outbound allow-list,
write scopes → read-only, then none
4 revoke credentials, OAuth grants, service
accounts, signing keys, session tokens
5 mark the artifacts provenance boundary + dates (STEP 4)
6 retire the record registry entry → decommissioned, with
the evidence, not deleted
Step 1 is the one that gets missed, and it is the one that produces zombies. An agent with no credentials but a live schedule still wakes up, still fails, still pages somebody, and still looks alive in your registry six months later. Enumerate the ingress exhaustively — a schedule is obvious, a webhook registered once by a departed engineer is not, and an event subscription that refreshes itself will quietly outlive everything else you turned off. Scheduled and triggered agents is the inventory of places to look.
Note that step 6 is not a delete. Removing the registry entry destroys the only index from the agent's name to its credentials, its tools and its outputs — which is the thing you will need when one of those outputs is questioned.
Draining is a decision about side effects, not a wait.
"Let it finish" is fine for a request-shaped service and meaningless for an agent, because the unit of work is a multi-step task that has already half-happened. A task that has created the purchase order and not yet sent the confirmation is in a state your shutdown does not understand.
So decide per task class, before you start, and write it down:
- Finish. For tasks whose remaining steps are bounded and whose partial state is worse than their completed state. Needs the credentials to stay live through the drain, which is why revocation is step 4 and not step 1.
- Abort and compensate. For tasks where you can reverse what happened — the compensating-action machinery in repairing agent side effects. Record the compensation, because an unlogged reversal is indistinguishable from a second error.
- Abort and escalate. For tasks you cannot reverse and cannot complete. These become a human's queue item with the trace attached. There will be more of these than you expect, and the number is the real measure of how reversible your agent was.
If the loop is event-sourced, the journal tells you exactly where each task stopped and what it had already done — the payoff described in durable state and resumability, and the reason a decommission is a good forcing function for finding out whether you have one. If it is not, your drain is a guess, and you should budget for the escalation pile.
The half with no off switch: four artifact classes, and the boundary you owe them.
Everything above is enumerable. This is the part that is not, and it is where the real liability sits. An agent that ran for a year left four kinds of residue that keep working after it stops:
- Records in systems of record. Tickets closed, invoices approved, CRM fields set, rows written. These are now indistinguishable from human-made records unless something marked them, which is the whole argument of disclosure and content provenance. If they were not marked at write time, you cannot retrofit the mark — but you can record the window and the scope.
- Memory and context other agents read. Facts it wrote into a shared store do not stop being retrieved because the writer is gone. Worse, they are now unattributed assertions with no author to question — see what a subagent inherits for how far they travel, and erasure against agent memory for what removal actually costs.
- Derived data. Indexes it populated, embeddings it generated, summaries that became the source for later summaries. Deleting the agent does not invalidate the index, and rebuilding it is a project — reindexing and embedding migrations.
- Documents humans now rely on. The runbook it drafted, the analysis in a deck, the policy summary someone pasted into a wiki. These have left your systems entirely and the only available control is telling people.
You cannot switch these off, so the deliverable is a provenance boundary: a dated statement of what this agent produced, where, under which model and policy version, from when to when. Three fields and a scope. It is cheap to write on the day you decommission and impossible to reconstruct a year later, and it is what converts "we think that was the old agent" into an answer — the evidentiary standard in audit trails, and the precondition for honouring contestability and appeals against a decision made by something that no longer exists.
Find the consumers before they find you.
The outage you cause by decommissioning cleanly is the one nobody modelled, because the dependency was never declared. An agent's consumers accumulate informally: another agent that reads its output, a dashboard whose numbers come from its writes, a weekly job that expects a file, a human whose morning starts with its digest.
Derive the list the same way you derive the inventory — from what the agent actually touched rather than from what anyone declared. Three queries cover most of it: who read the tables it wrote in the last 90 days; which agents retrieved from the memory namespace it wrote to; and which tool-call traces name it as an upstream. Then announce a date, not an intention, and keep the agent read-only through the notice period rather than dark — a read-only agent that answers from its existing state is a far gentler deprecation than a 404, and it is the same courtesy you would extend under model deprecation and migration.
Watch for the consumer that is another agent, because it will not complain — it will route around the gap, invent a value, or fail silently in a way that shows up as a quality regression three weeks later with no obvious cause. Make agent-to-agent dependencies an explicit field in the registry, or accept that you will discover them this way.
Finish it with a record and a resurrection test.
A decommission is done when someone can prove it happened and nobody can quietly undo it. Those are two different properties and most teams get neither.
- The record is the deliverable. Registry entry marked decommissioned with a date, the provenance boundary from step 4, the list of revoked grants with their revocation timestamps, the escalation pile from step 3, and the named owner who signed it off. This is what accountability and roles means when the subject no longer exists to be asked.
- Test the resurrection path. Try to turn it back on. If flipping one feature flag brings it back with live credentials, you have paused an agent and told your auditor you retired one. A real decommission requires re-issuing something — a credential, a grant, a registry entry — so that the comeback is an event somebody authorises.
- Remove it from the tool catalog in both directions. Its own tools go, and so does its entry anywhere other agents could discover and call it. A stale catalog entry is how a decommissioned agent gets reinvoked by something that never knew it was gone — tool catalog lifecycle.
- Keep what retention requires, delete the rest deliberately. Traces and decision records usually have a retention obligation that outlives the agent by years; working state usually does not. Decide which is which on the day, under retention and legal hold, rather than letting the default be "keep everything forever in a bucket nobody owns".
Do this today: pick the agent in your estate that is least obviously still needed and run just steps 1 and 4 on paper — list every ingress that could wake it, and write the four-line provenance boundary for what it has produced. Most teams discover an ingress they had forgotten within the first ten minutes, and the provenance boundary takes about as long as this page took to read. Then make the decommissioning record a required field on the registry entry, so the 21% number stops describing you — because the cheapest moment to decide how an agent will be retired is the day you onboard it.