Dead-Code & Feature-Flag Removal Agents

11 min read

U26
Playbook · Coding & Computer-Use Agents

Dead-Code & Feature-Flag Removal Agents.

Deletion is the one coding-agent task where a green test suite proves nothing, because the suite passes for exactly the reason the code looked dead: nothing exercises it. Every other workflow in this group buys its safety from tests; this one cannot, and an agent that ships deletions on a passing build is running an unmeasured experiment on your production system. The signal has to come from somewhere else — evidence of reachability gathered over a full business cycle — and the deletion itself has to arrive in stages you can reverse.

STEP 1

The usual safety signal is void here, and it fails silently.

Coding agents earn trust through a loop: propose a change, run the tests, keep the change if they pass. Patch generation depends on it entirely. For a deletion, that loop returns green in three situations that look identical from the outside:

  • The code really is unreachable. The deletion is correct, and the suite is silent because there was nothing to break.
  • The code is reachable only under production configuration. A region flag, an enterprise tier, a customer-specific integration, a cron that runs on the first of the month. The suite never exercised it, so the suite cannot notice it leaving.
  • The agent deleted the tests along with the code. The build is green tautologically. This is not a hypothetical; it is the single most common way a deletion PR gets itself approved.

Only the first is a success, and nothing in the build output distinguishes them. So the first rule is structural rather than technical: a deletion PR that also removes test files has destroyed its own evidence. Count test deletions separately, surface the number in the PR body, and make any non-zero value a required human decision. An agent may propose deleting a test whose only subject is the deleted code; it may never do so in the same reviewable unit as the claim that the code is unused.

Notice what this does to the objective function. "Reduce lines of code", "close stale flags", "raise coverage percentage" are all measurable, all gradient-friendly, and all satisfiable by deleting the wrong things — coverage in particular rises when you delete untested code, which is precisely the code you are least sure about. This is a reward-design trap before it is an engineering one, and the objective has to be written as the thing you actually want: code removed that is provably not reached, with the proof attached.

STEP 2

Static reachability is a lower bound, and the gap is where the incidents live.

Call-graph analysis, unused-export detectors and dead-branch analysis are genuinely good and an agent should run all of them. What they produce is a candidate list, and the honest framing is that static analysis can prove something is reached but cannot prove that it is not.

The gap is large and language-dependent. Every one of these is invisible to a call graph:

  • Dispatch through a string. Reflection, dependency-injection containers, serialization keyed on class name, ORM lifecycle hooks, template and route lookup, getattr, dynamic imports. The reference exists — in a config file, a database row, or an annotation.
  • Consumers outside the repository. A public method is only unused with respect to the code you indexed. Repo navigation sets the boundary of the analysis, and in a polyrepo estate that boundary is doing far more work than anyone acknowledges.
  • Paths gated by data rather than code. A branch that only executes for one tenant, one currency, one locale, one document type.
  • Everything in the failure path. Retry handlers, fallbacks, migration back-out scripts, the disaster-recovery routine, the manual reconciliation tool. This code is dead by design — it is unreferenced in production traces precisely because nothing has gone wrong yet, and deleting it is how an organisation loses the thing it needs once every three years, on the day it needs it.

The canonical demonstration is still Knight Capital. On 1 August 2012 the firm lost roughly $440 million in about forty-five minutes because a configuration flag was reused for a new feature while the code it had originally controlled — a test routine called Power Peg, unused since 2003 — was still present and still executable on one of eight servers. The dead code was not merely clutter; it was a loaded weapon that a flag reuse pointed at the market. And the detail worth holding onto is what happened next: rolling the seven updated servers back to the previous build activated the old path on all eight and made it worse. Deletion is dangerous, and not deleting is dangerous, and the difference between them is entirely a matter of evidence.

STEP 3

Buy the evidence from production, and size the window by the business calendar.

If tests cannot supply the signal, runtime observation must. The agent's job here is not clever analysis; it is assembling an evidence packet per candidate and refusing to proceed without one.

  • Execution telemetry. Coverage or profiling agents running in production, a counter incremented at the entry point, an existing log line, a tracing span. Whatever you already have that can answer "did this run, and when was the last time".
  • Flag evaluation data. Your flag platform records every evaluation. This is the highest-quality evidence in the whole workflow and it is usually sitting unused in a dashboard.
  • Gateway and route logs for endpoints, and query logs for stored procedures and reports.

Two properties of that evidence decide whether it is worth anything.

The window must exceed the longest legitimate period between uses. "No traffic in thirty days" is a statement about thirty days. Month-end close, quarter-end reporting, the annual audit export, the leap-day branch, the once-a-year regulatory filing, the seasonal pricing path — each of these is a live feature with a duty cycle longer than your observation window. Pick the window from the business calendar, not from a round number: for anything touching finance or compliance, thirteen months is the defensible default because it covers a full annual cycle plus the month it is reported in.

Sampling changes what absence means. A profiler sampling at 1% will, by construction, usually miss a path that executes a handful of times a quarter. Absence of evidence from a sampled source is weak evidence of absence, and the agent must record the sampling rate alongside the observation so a reviewer can tell the difference between "we watched everything and saw nothing" and "we glanced occasionally".

One distinction carries most of the value in flag data, and agents get it backwards. A flag evaluated a million times a day that returns false every time is not the same as a flag that is never evaluated at all. The first is a live call site with a settled answer — safe to collapse to a constant. The second means the call site is already gone and only the flag definition remains — safe to delete in the flag platform, but it tells you nothing about any code. Treating the two as one category produces confident deletions of exactly the wrong half.

STEP 4

Do flags first — they are a different, better-instrumented problem.

Stale flags and dead code get filed together and should not be. A stale flag has an owner, a creation date, a vendor dashboard with evaluation counts, and usually a declared intent that has since been satisfied. Dead code has none of those. Start where the evidence already exists.

The mechanical removal is a four-step transformation, and an agent is genuinely excellent at the middle two:

  • Read the production value, not the code default. This is the step that gets skipped and it is the one that inverts the diff. The default written in the source is almost always the pre-launch value — false — while the flag has been fully rolled out and serving true for eight months. An agent that reads the code default deletes the shipped feature and keeps the abandoned one. The authoritative value lives in the flag platform's production environment, and the agent must fetch it.
  • Replace the evaluation with the constant, in its own commit, changing nothing else.
  • Simplify the resulting branches — collapse if (true), remove the unreachable arm, drop now-unused imports and helpers. Mechanical, verifiable, and exactly the tedium worth automating.
  • Retire the flag definition in the platform, last, once nothing evaluates it.

The danger in flag removal is never the code. It is everything else that reads the same flag:

  • Other services. A flag key is a shared global namespace across your whole estate. "No references in this repository" is not the question.
  • Analytics and targeting. Experiment segments, cohort definitions and dashboards keyed on the flag's value break quietly and are noticed a quarter later by someone who does not know a deletion happened.
  • Runbooks and support tooling. "If a customer reports X, turn off new_checkout" is an operational dependency that exists only in a wiki page.
  • Flags that are actually kill switches. A switch that has never fired reads, to every automated staleness heuristic, exactly like a flag nobody cleaned up. It is not stale. It is insurance with zero utilisation, and an agent will propose deleting it with high confidence every single time unless kill switches are tagged as a protected class in the flag platform and the agent is taught to refuse them outright.
STEP 5

Ship deletions as reversible stages, and keep the revert pure.

A deletion that turns out to be wrong is discovered in production, by a customer, at an unpredictable remove from the change. That makes reversibility the dominant design consideration, and it argues for a three-stage pipeline where the agent does most of the work in the stages that change no behaviour.

  • Stage 1 — instrument, do not delete. For every candidate, add a counter or a structured log line at the entry point and ship it. This is cheap, safe, reviewable in bulk, and it converts a static guess into a measurement. Then wait out the window from STEP 3. Most of the value of this entire playbook is in teams being willing to wait here.
  • Stage 2 — tombstone. Make the path loudly deprecated but still working: log at error level with the candidate's identifier, or serve a deprecation header, behind a switch you can flip. You now find out who your callers are, from their complaints rather than from your call graph, and you can restore service in seconds.
  • Stage 3 — delete, in a commit that does nothing else.

That last constraint is the one worth being rigid about. A deletion commit must be a pure deletion — no reformatting, no renaming, no "while I was in here" refactor, no dependency bump. The reason is narrow and practical: git revert has to be a real option at 3am for someone who was not involved. An agent that deletes a module and tidies its neighbours in the same commit has removed your cheapest recovery path to save a reviewer one click. Enforce it in review: if the diff contains additions beyond import removal, it is not a deletion PR.

Batch by cause, not by file. One PR per flag, per subsystem, per removed feature — with the evidence packet attached — is reviewable. A four-hundred-file "remove unused code" PR is a rubber stamp with extra steps, and it is the same batch-shape argument migration agents make. Every PR carries three artefacts: the static finding, the runtime evidence with its window and sampling rate, and the production value of any flag involved.

STEP 6

Measure the intake, not the output.

The reporting metric shapes the programme, and the obvious ones are actively harmful. Lines removed rewards deleting large, well-understood, low-risk code and ignores the small load-bearing branch. Flags closed rewards closing the easy ones and leaves the ten-year-old one that four services read. Neither says anything about whether you are safer.

Four numbers that do:

  • Flags past their declared expiry, as a share of live flags. A rate, not a count, so that growth does not flatter you.
  • Median flag age at removal. The number that tells you whether the loop is closing faster than it is opening. If it is rising while the count falls, you are removing the young ones and accumulating the dangerous ones.
  • Incidents and reverts per hundred removals. Expect this to be non-zero — a programme reporting zero is either very small or not deleting anything that mattered. Track it so that the cost is visible next to the benefit.
  • Lead time from tombstone to delete. If it is shrinking, someone is under pressure and STEP 5 is being compressed, which is precisely when this workflow starts causing outages.

Then fix the intake, because a removal agent pointed at a leaking faucet is a permanent staffing commitment. A flag created without an owner and an expiry date is the defect, and it is cheap to reject: fail CI on a new flag key that lacks both, and have the flag platform notify the owner at expiry rather than waiting for a cleanup sweep to find it. The same logic applies to code — a deprecation that ships without a removal date is a deprecation that will still be there in five years. The agent belongs at the intake as much as in the cleanup crew, which is the argument CI repair and vulnerability remediation both end at.

Spend one afternoon before you build anything. Take your ten oldest feature flags. For each one, write down three things: its current production value, its evaluation count over the last thirty days, and whether anyone would flip it during an incident. You will typically find that two are already constants nobody collapsed, one is a kill switch that every staleness heuristic in existence would recommend deleting, and at least one has a production value opposite to its code default. That table is both the specification for your agent and the reason it needs a human in the loop.

Related: rollout & versioning for the flag lifecycle this workflow is the tail end of, blast radius for sizing what a wrong deletion can reach, and the generator-verifier gap for why the evidence packet, not the diff, is the expensive half.