Access Reviews for Agent Credentials

7 min read

C23
Operation · Governance & Compliance

Access reviews were built around a termination date. Your agents don't have one.

Every access review programme in your organisation works because HR emits an event when a person leaves, and that event is what actually cleans up entitlements — the quarterly attestation is the paperwork on top of it. An agent has no leaver event. Its credentials are created by an engineer in the middle of a project, they outlive the project, the engineer, and usually the use case, and nothing anywhere signals that they should stop. The fix is not a bigger spreadsheet with service principals added to it: it is to stop reviewing credentials at all and start reviewing reachable calls, with expiry driven by use rather than by a reviewer's memory.

STEP 1

Name why the human lifecycle does not transfer.

Joiner–mover–leaver is a good design and it rests on three properties that agents do not have. Say them out loud before you try to reuse the process, because each one breaks a different control.

  • No termination event. A person's access decays because someone else's system announces the end. An agent's access ends when a human remembers, which is to say it does not end.
  • No manager who can attest. Certification assumes a reviewer who knows what the subject does all day. Ask a director to attest that svc-agent-prod-07 still needs write access to the billing database and you will get an approval, because the alternative is breaking something on a Friday. Rubber-stamping is not a discipline failure here; it is the rational response to an unreadable question.
  • No stable job function. A human's entitlements drift slowly with role changes. An agent's effective reach changes the moment someone edits a prompt or registers a tool — a deploy, not a transfer. Your review cadence is quarterly and your change rate is daily.

The practical consequence is that the finding an auditor writes — "service accounts are not included in the periodic access review" — has the wrong remedy attached to it. Adding them to the same review produces attestation at scale and cleanup of nothing. What follows is the remedy that works, and it happens to produce better evidence than the attestation did.

STEP 2

Review the reachable call, not the credential row.

The unit matters more than the cadence. A credential inventory undercounts an agent's reach in both directions, and the gap is where incidents live.

  • One credential, many capabilities. A single database role, a single cloud role, a single OAuth grant with six scopes. Reviewing the row tells you nothing about which of the six the agent uses.
  • Many credentials, one agent. The agent's own key, the MCP server's separate key, the connector platform's per-user token, the shared read-only analytics account somebody reused. Different owners, different systems, one blast radius.
  • Reach nobody issued. A browser profile with live sessions, a network path with no egress policy, a filesystem mount. This is ambient authority, and it appears in no credential store at all.

So the reviewable object is this agent can perform this call against this system on this data. Derive it from the tool registry joined to what each credential actually grants, and keep it next to the agent inventory rather than in the IAM tool — the IAM tool cannot see your tool definitions.

// The row a reviewer can actually answer.
{ "agent": "invoice-triage",
  "capability": "erp.invoice.approve",
  "credential": "svc-invoice-triage",
  "limit": { "max_amount_eur": 5000, "rate_per_hour": 40 },
  "granted_for": "JIRA-4412 — AP automation pilot",
  "owner": "ap-platform-team",
  "last_successful_use": "2026-04-02T09:14:22Z",
  "expires": "2026-10-01" }
STEP 3

Make each grant reviewable in one sentence, or it will be approved unread.

A reviewer can only decline what they can understand. Three fields turn an unanswerable row into an answerable one, and all three must be captured at grant time — none of them can be reconstructed later.

  • The reason, as a link. Not "production access" but the ticket, incident or project that justified it. A grant whose reason is a closed ticket answers its own review question.
  • The named owner. A team, not an individual, because individuals leave and the grant does not. This is the same role the accountability page assigns to the operator.
  • The bound limit. Amount ceilings, rate caps, row-level scope. A capability with a limit attached is reviewable on the limit — "should this still be €5,000?" is a question a business owner can answer, and "should this agent have ERP access?" is not.

Write the reason in the language of the business process, not the infrastructure. The test is whether someone outside the platform team can read the row and form an opinion; if they cannot, the review is theatre regardless of how often you run it.

STEP 4

Replace the attestation with use-based expiry.

This is the control that actually removes access, and the data for it already exists in systems you pay for. AWS IAM reports last-accessed information per service and per action; cloud identity providers log sign-ins for workload identities; your own gateway logs every tool call. The rule is mechanical.

  • Every capability gets an expiry at grant time. Ninety days is a reasonable default for a pilot, a year for a capability that has survived one full review. No capability is granted without one, and "permanent" is not an option on the form.
  • Unused for N days, it is revoked. Not flagged, not escalated — revoked, with a notification to the owner and a one-click restore for a week. An unused grant is the cheapest possible thing to remove and the hardest possible thing to justify keeping.
  • Used but never near the limit, it shrinks. If the approval ceiling is €5,000 and the p99 actual is €180, the ceiling is wrong. Ratchet limits down toward observed use on a schedule; this is where an access review turns into a real reduction in blast radius instead of a list that got re-approved.
  • Re-grant is cheap or the rule will be bypassed. If restoring a revoked capability takes three days and two approvals, engineers will keep everything alive by scheduling a monthly call that touches it. Make the restore path fast enough that hoarding has no payoff.

Pair this with genuinely short-lived credentials underneath, so that what you are reviewing is the right to mint, not a secret sitting in a vault — see scoped credentials for agents.

STEP 5

The delegated half does have a leaver event. Wire it.

Some of an agent's access is not the agent's at all — it holds a token that represents a specific human who clicked consent. Mailbox access, a calendar, a CRM seat, a payment mandate. Here the termination event exists, it is the one HR already emits, and it is almost never connected to the agent platform.

  • Index delegated grants by the human, not by the agent. "What did Dana authorise?" has to be a query, because that is the question that arrives on Dana's last day and also the question that arrives with an erasure request.
  • Revoke on the leaver event, automatically. An agent still acting on a departed employee's mandate is the cleanest audit finding there is, and it is entirely mechanical to prevent.
  • Expire consent independently of the token. An OAuth refresh token does not expire because the business reason did. Keep your own consent record with its own end date and treat the provider's token as an implementation detail.
  • Distinguish the two identities in every log line. Acting agent and authorising human are different fields. Collapsing them is what makes an audit trail unable to answer who was responsible.
STEP 6

Produce evidence that survives an auditor, not a screenshot.

The question you will be asked is not "do you run access reviews" but "show me that this agent's access matched its purpose throughout the period". Four artefacts answer it, and each is a query rather than a document.

  • The capability register, versioned — what existed, when, with which limits. Point-in-time reconstruction is the whole requirement; a current-state export proves nothing about March.
  • The revocation log. A review that never removes anything is evidence of a broken review. Removals per period is the single best health metric for this programme, and a flat zero should worry you more than a large number.
  • Exceptions, with expiry. Every programme has standing exceptions; unbounded ones are how a temporary elevation becomes permanent. An exception without an end date is a finding you wrote yourself.
  • Coverage. The share of live agent capabilities that are in the register at all. Discovered-but-unregistered is the number that tells you whether the register describes reality — the same distinction inventory draws between discovery and an inventory.

Start with one query, this week: list every credential your agents can use that has had no successful call in ninety days. In most organisations that list is long, nobody is defending any row on it, and revoking it is a measurable reduction in reachable surface with essentially no risk of breaking production. Do that once manually, then make it a scheduled job — you will have built the only part of an access review that ever actually removed access, and the register and the attestation can be assembled around it afterwards. Related: detecting agent compromise for what to watch on the grants you keep, and secrets management for what sits under them.