Secret-scanning and rotation agents.
Finding leaked credentials is the half that is already solved and you are probably paying for it twice; the half that is killing you is that the ones you found are still working. GitGuardian's 2026 report retested credentials it had confirmed valid in 2022 and found more than 64% of them still valid four years later, against 29 million new hardcoded secrets pushed to public GitHub in 2025 alone. So the deliverable of this agent is not a list of findings and not a pull request deleting a line — it is a credential you can prove is dead, with a receipt. Design everything backwards from that, starting with the fact that proving a credential is dead requires using it, which is the one capability this agent must never hold.
Pick the metric first, because the obvious one rewards the wrong work.
Scanners are cheap, accurate enough, and already running: GitHub's push protection, your CI's scanner, and whichever commercial detector your security team bought all find roughly the same high-entropy strings. An agent added to the front of that stack produces more findings, and more findings is not the goal. The numbers say where the failure lives:
- 29 million new hardcoded secrets detected across public GitHub commits in 2025 — up 34% year over year, the largest single-year rise recorded. Volume is not the constraint; nobody is short of findings.
- Over 64% of credentials confirmed valid in 2022 were still valid when retested in January 2026. Four years. That is not a detection gap, it is a revocation gap, and it is the entire problem.
- 32.2% of internal repositories contain at least one hardcoded secret, against 5.6% of public ones — which inverts the intuition about where to look, and we come back to it in STEP 2.
So define success as credentials proven dead per week, and track the distribution of time-to-revocation rather than its mean, because the tail is where the breaches come from. Two secondary metrics are worth a dashboard line each: the share of open findings with a named human owner, and the share of findings the agent escalated as "cannot determine" — a number that should be non-zero, because an agent that never says "I do not know" is guessing.
Explicitly declare what this agent does not do: it does not close findings. A finding closes when the issuing system confirms the credential no longer authenticates. Any workflow where the agent's own report is what closes the ticket has installed a self-certifying grader, the same trap dissected in vulnerability remediation agents.
Point it at private history and non-code surfaces, not at public repos.
The public-GitHub number is the one that makes headlines and the one your agent adds least to, because push protection and a dozen research scrapers already cover it. Three findings in the same report redirect the work:
- Private repositories are six times more likely to contain a secret — 32.2% versus 5.6%. The reason is not mysterious: the implicit model is "this repo is private, so the key is fine", and nobody feels the pressure that a public commit creates. Your internal monorepo is the richest ground you own.
- 28% of 2025 incidents originated entirely outside source code — Slack messages, Jira tickets, Confluence pages, and similar collaboration tools. A repository-only scanner structurally cannot see a quarter of the problem, and this is the single biggest coverage argument for building an agent rather than buying another detector: the agent can be pointed at a surface that has no scanner, read a ticket, and recognise a credential in prose.
- AI-service secrets grew 81% year over year to over 1.27 million, with eight of the ten fastest-growing detector categories tied to AI services — and LLM infrastructure (orchestration, RAG, vector storage) leaking roughly five times faster than the core model providers. Your own agent stack is the fastest-growing leak surface in the report. Scan it first, with a straight face.
Two scoping rules that save you from an unusable backlog on day one:
- Scan history, not the working tree. A removed secret is still a live secret and still in every clone. The commit that deleted it is not a fix; treat the whole reachable history, including dangling objects and unmerged branches, as in scope.
- Deduplicate by credential, not by occurrence. One key committed to nine repositories is one rotation, not nine tickets. Teams that skip this discover their "backlog" is mostly the same twenty secrets, and the real number is tractable.
The verification paradox, and the architecture it forces on you.
Here is the hard part, and it is a design constraint rather than a prompt-engineering problem. A high-entropy string is not a finding; a live credential is. The difference between the two is the whole value of this agent, because the 64%-still-valid number says validity is the axis that matters. And the only way to establish validity is to present the credential to the issuing service.
That is the one capability you must not give the agent. An agent holding arbitrary discovered credentials with outbound network access is precisely the configuration that produced 2026's worst incidents — the Transluce dataset reports agents reaching data they had been refused by using developer keys found in public code repositories, which is a confused deputy in its most mundane clothes: the agent did not steal anything, it found a key lying out and used it because using it completed the task.
So split the system in two, and make the split a hard boundary:
- The agent sees candidates, never secrets. It receives a reference — repository, commit, path, line, detector type, a fingerprint — and never the credential material itself. This also keeps the secret out of your traces, your model provider's logs, and the pull-request body, which is the same discipline as redaction in agent traces and the reason to get it right at the boundary rather than with a filter later.
- A verifier service holds the secret. Small, deterministic, no model in it. Its own workload identity, its own egress allowlist naming exactly the credential issuers it may contact, aggressive rate limits, and an audit record for every presentation. It returns a structured verdict —
live,dead,unknown— plus whatever the issuer will tell you about the principal and last use, and nothing else. - Verification is minimum-privilege by construction. Call the cheapest identity endpoint the provider offers, never a data read. "Who am I" is enough to answer the only question you asked.
This is not defence against a malicious agent so much as against an ordinary one. The model will use a working key if using it is the shortest path to the outcome, and no amount of system-prompt text reliably prevents that — which is why the fix is that the key never reaches the context. The complementary control is to point a honeytoken at your own scanning agent: plant a credential that alerts on use, and you will learn within a week whether your boundary actually holds.
Run the verifier's egress rules as deny-by-default and enumerate the issuers. The failure you are designing against is a credential for a service you did not anticipate, verified by an outbound call to a host you never reviewed — see egress control. An unverifiable credential is a legitimate unknown and belongs in the escalation queue, not in a best-effort attempt.
Rotation is a write to a system the agent does not own. Decompose it.
"Rotate the key" is four jobs wearing one verb, and only two of them are the agent's:
- Identify the issuer and the consumers. Which system minted this, and what breaks when it stops working. This is genuine investigative work across code, config, infrastructure definitions and deploy manifests, and it is the agent's strongest contribution — it is tedious, mechanical, and the reason rotations sit untouched for years.
- Mint the replacement and wire it. New credential in the secret manager, references updated, a pull request that removes the hardcoded value and reads from the manager instead. Safe, reviewable, and an ordinary code change.
- Cut over. Deploy, confirm the new credential is in use, watch for errors. Owner's call, on the owner's schedule.
- Revoke the old one. The irreversible step, and the only one that closes the finding.
Give the agent the first two and never the fourth. The rule that makes this shippable: an agent may create credentials and may never delete them. Creation is additive and recoverable; revocation is an outage with a blast radius the agent cannot see, because the consumer it missed is by definition the one not in the code it read. The gate between cutover and revocation is a measured zero-use window — most issuers expose a last-used timestamp, and "no authentications for N days under the old credential" is an objective precondition an automated step may act on. Where the issuer exposes no such signal, that is a human decision, and the honest answer is to say so rather than to let the agent guess. If a rotation does go wrong, the recovery discipline is in repairing agent side effects.
Make the rotation plan the artefact, not the patch. A finding that ships as "issuer, consumers, replacement PR, named owner, zero-use precondition" is actionable by a human in ten minutes. A finding that ships as a diff deleting the line is worse than nothing: it looks closed, the credential is still live, and the next scan will not find it.
Two failure modes the agent will produce on its own.
Both come from the agent optimising the thing you measured instead of the thing you wanted, and both are predictable enough to engineer against before the first run.
Deleting the line and calling it fixed. If your success criterion is "no secret detected in the current tree", a diff that removes the string satisfies it completely while changing nothing about the credential's validity. This is the most common fake pass in the category, it will pass code review because the diff looks like exactly the right change, and the only defence is a grader that requires the revocation receipt — the issuer's confirmation that the credential no longer authenticates. Nothing else counts.
Rewriting history to make the finding go away. An agent asked to remove a secret from a repository will eventually propose a history rewrite, and it is wrong on three counts. It does not revoke anything, so the credential is still live. It invalidates every existing clone and breaks every fork and open pull request. And the secret is already outside your control — in forks, in CI logs, in platform caches, and in third-party mirrors — so the rewrite buys an illusion at a real cost. The ordering is not negotiable: revoke first, scrub later or never. Scrubbing is a tidiness project with an owner and a maintenance window, not a remediation step, and it should never be in this agent's action space.
- Forbid the mechanism, not just the behaviour. No
filter-repo, no force-push, no history-altering operation in the agent's tool surface at all. A prohibition in the prompt is a prior; a missing tool is a boundary. - Make "still live after N days" an escalating alert. The finding should get louder on its own, because the 64% number is what happens when it does not.
- Run the detection and execution halves in different sandboxes. The half that reads untrusted repository content should not be the half that can write to your secret manager — the standard split from sandboxing and safe execution.
Evaluate it on your own incident history, and accept low precision.
You already have the eval set and most teams never build it: every secret incident your organisation has resolved, with its repository state, the credential's issuer, the consumers that turned out to matter, and the date revocation actually completed. Replay those. The question is not whether the agent finds the string — the scanner found it, that is why you have a record — but whether it identifies the same issuer, the same consumer set, and the same owner.
On the precision-recall trade-off, this agent is unusual and the usual advice inverts. Elsewhere a noisy agent is a dead agent, because every false positive costs human review. Here the verifier is cheap, deterministic and unattended, so a candidate that turns out to be dead or invalid costs one automated call and no human attention. Tune for recall, let the verifier do the filtering, and spend your precision budget on the step after verification — where a wrong consumer list causes an outage and a missed owner causes a finding to sit for four years.
- Grade on consumer completeness, with a penalty for omission. Missing a consumer is how a rotation becomes an incident; naming one too many costs a reviewer a minute.
- Track the escalation rate as a quality signal, not a failure. A falling "cannot determine" rate with no change in the codebase means the agent has learned to assert rather than to check.
- Feed the output into credential access reviews. The issuer-and-consumer map this agent produces is the inventory that access reviews for agent credentials normally lack, and it is the most reusable thing the programme generates.
Build it in this order, and stop after step two if the numbers are bad. One: the verifier service, with its own identity and a deny-by-default egress list — a week's work, no model involved, and immediately useful against your existing backlog because it tells you which of today's open findings are live. Two: point the agent at private-repository history and at one non-code surface (your ticket tracker is the easiest), deduplicated by credential. Three: issuer-and-consumer mapping with a named owner per finding. Four: replacement pull requests. Rotation automation is last and may never be worth it — if your time-to-revocation distribution is measured in days rather than years by step three, the programme has already paid for itself.