Separation of Duties for Agents

9 min read

S9
Deep Dive · Agent Security

Two signatures, one failure.

Maker–checker buys you almost nothing against the threat you installed it for: the document that fools the maker is the document that fools the checker, so two-of-two approval on an injected input is one approval wearing two hats. Separation of duties was never a redundancy argument — it is an independence argument, and an agent stack destroys independence by default, quietly, while the audit log fills with the right number of entries. The number that tells you which one you have is the disagreement rate, and it is the one number nobody puts on a dashboard.

STEP 1

The control is a claim about independence, not about counting.

Separation of duties is old and it is load-bearing in places where money moves: the person who raises a payment may not be the person who releases it, the developer who writes the change may not be the one who deploys it to production. The control is usually described as "two people must agree", which is the ceremony rather than the mechanism. What makes it work is that the second person fails for different reasons than the first.

Three properties do that work, and classical SoD gets all three for free because the two parties are different humans in different roles:

  • Uncorrelated error. The checker's mistakes are not the maker's mistakes. A tired requester and a careless approver fail at different times, on different invoices, for different reasons.
  • Different information. The approver holds something the requester does not — the budget, the vendor history, the knowledge that this supplier was terminated last quarter. Review adds a fact, not just a glance.
  • Different incentives. Collusion requires two people to decide to collude, and each has something to lose. The cost of defeating the control is a conspiracy rather than a mistake.

Now replace the checker with an agent. The ceremony survives intact — there are two steps, two log lines, two identities in the receipt. Every one of the three properties is now an open question, and the default answers are bad. This is the same failure the reference monitor test catches in a different guise: a control that the thing being controlled can influence is not a control, and a second agent reading the same context is influenced by exactly what influenced the first.

The uncomfortable version: if your checker is a second call to the same model with a "review this carefully" prompt, you have not added a control. You have added a sampling step, and the ensemble literature is explicit that correlated samples buy sub-linear improvement — see agent debate and ensembles, where diversity is the load-bearing variable. A control that behaves like an ensemble should be budgeted like one.

STEP 2

Four couplings, and you probably have three of them.

Independence does not fail in one way. It fails along four axes that are easy to audit individually and almost never audited together, because each one lives with a different team.

  • Shared model. Same weights, or same family, or merely trained on overlapping data. Self-preference is measurable: an LLM asked to judge candidate outputs favours its own generations, which means the maker's output is systematically easier to approve than a stranger's. That is the opposite of what a checker is for — see judge calibration and meta-evaluation.
  • Shared context. The checker is handed the maker's transcript, plan or summary. Everything the maker read, the checker now reads, including the untrusted part. This is the coupling that matters most and the one that feels most like diligence.
  • Shared authority. One credential, one network path, one browser profile. If the maker can perform the action without the checker, the checker is advisory; if the checker's approval is just a boolean the maker's own loop sets, the second signature is being produced by the first party. Ambient authority is where this hides.
  • Shared objective. Both agents are scored on completing the task. A checker whose implicit goal is "unblock the workflow" converges on approval, exactly as human reviewers do when the alternative is breaking something on a Friday.
# Score your own maker/checker pair. Any ✗ is a coupling you are carrying.

                     model    context   authority  objective
same call, "review"    ✗         ✗          ✗          ✗
2nd call, same model   ✗         ✗          ~          ✗
2nd model, same ctx    ✓         ✗          ~          ✗
2nd model, task only   ✓         ✓          ~          ✗
policy engine + rules  ✓         ✓          ✓          ✓
human reviewer         ✓         ~          ✓          ~

✓ independent   ~ partially   ✗ coupled

Read the middle rows carefully, because that is where most production systems sit. Swapping in a second vendor's model fixes one column and leaves three. The row people believe they have built is the last one; the row they have built is usually the second or third.

STEP 3

The arithmetic, and the inversion that follows from it.

Put numbers on it, because the intuition is wrong in a specific direction. Let the maker propose a bad action with probability p. Let the checker, when independent, catch a bad action with probability q. Let c be the probability that the checker fails for the same reason the maker did — it read the same poisoned page, it shares the blind spot, it recognises its own output as correct.

# escape rate = p × (1 − (1 − c) × q)        p = 0.05, q = 0.80

c = 0.0   escape 1.0%    risk ÷ 5.0     the control you costed
c = 0.3   escape 1.8%    risk ÷ 2.8
c = 0.5   escape 3.0%    risk ÷ 1.7
c = 0.9   escape 4.6%    risk ÷ 1.1     the control you have
c = 1.0   escape 5.0%    risk ÷ 1.0     ceremony

Two things follow. The first is that c dominates q: a mediocre but independent checker beats an excellent correlated one, so effort spent making the checker smarter is usually worse than effort spent making it different. The second is the inversion, and it is the whole reason this page exists.

c is not a constant. It is near zero for the maker's random slips — a transposed account number, a hallucinated field name — and near one for precisely the failures you were afraid of. An injected instruction in a retrieved document is read by both parties and is persuasive to both. A tool description that misstates what the tool does misleads both. A goal that drifted during a long run has drifted in the transcript the checker is reading. So the control's effectiveness is anti-correlated with the severity of the threat: it works on typos and evaporates on adversarial input.

This is why "we have a reviewer agent" is not an answer to prompt injection, and why the prompt-injection defense stack puts its weight on isolation and egress rather than on a second opinion. A checker that shares the maker's context is a detector of accidents, and you should budget it as one.

STEP 4

Disagreement rate is the metric, and it has an auditable floor.

You cannot measure c directly. You can measure how often the checker disagrees with the maker, and that number is bounded below by whatever you claim about the pair. If the maker is wrong 5% of the time and the checker catches 80% of those, and the checker also objects to 2% of good actions, you should observe disagreement around 0.05 × 0.8 + 0.95 × 0.02 ≈ 5.9%. Observe 0.3% instead and exactly one of your two beliefs is false: either the maker is far better than your error budget says, or the checker is not checking.

# The inequality an auditor can run in one query.

observed_disagreement  ≥  claimed_error_rate × claimed_catch_rate

# If it fails, one of the three numbers is wrong. Find out which
# before you find out in an incident.

Three refinements make the number trustworthy. Log the checker's verdict before the maker's action resolves, or you are measuring hindsight. Separate the two directions — blocked-and-was-bad against blocked-and-was-fine — because a checker tuned to near-zero false positives is a checker tuned to near-zero catches. And seed the stream deliberately: inject known-bad proposals at a low rate and confirm the checker rejects them, which is the only way to get a catch rate rather than an agreement rate. The same reasoning drives the 98%-approval argument in review queues for agent output, and the statistics for sizing the sample are in eval variance and statistical power.

Watch for the degenerate case that looks healthiest of all. A checker with 0% disagreement and a 100% approval rate produces a clean dashboard, a complete audit trail and no control whatsoever. Nothing in the record distinguishes it from a working one, which is why the floor above belongs in your monitoring and not in a quarterly review.

STEP 5

Buy independence on three axes, cheapest first.

Independence is purchasable, and the cheap purchases are not the ones teams reach for. Ranked by assurance per unit of effort:

  • Different information, not more of it. Give the checker the proposed action and its arguments — the resolved SQL, the payee and amount, the diff, the destination host — and nothing from the maker's transcript. A checker that cannot read the injected document cannot be injected by it. This single change moves c more than any model swap, and it makes the checker cheaper rather than more expensive.
  • Different mechanism where the invariant allows it. Most things you want a second opinion on are not opinions. "This payee appears in the approved-vendor table", "this diff touches no file under infra/", "this total matches the line items" are rules, and a rule has c = 0 against an adversary who cannot edit it. Spend the model on the residue — see policy-as-code for agents.
  • Different authority, enforced outside the maker. The checker's verdict must reach the effect through a path the maker does not control. If the approval is a field in the maker's own state, a retry or a resumed run will produce it again — which is how approvals get replayed rather than re-earned.
  • Different model, last. Useful, real, and the most expensive of the four. It addresses one column of the table in step 2 and tends to be treated as if it addressed all of them.

Two design notes that come up every time. First, give the checker a deny that terminates, not a comment: a verdict that becomes advice the maker may weigh has collapsed back to one party. Second, make the checker's output structured and enumerated rather than prose, so the reason a thing was blocked is machine-readable and comparable across runs — the argument in structured refusal and why-trails, and the input to the receipt in decision receipts and audit.

STEP 6

When you cannot get independence, stop counting the control.

Some checks genuinely need the maker's full context to be meaningful — "is this the right fix for the bug described in the ticket" cannot be evaluated from the diff alone. There, independence is not available at any price, and the honest move is to write that down rather than to buy a second model and feel better.

  • Name the coupling in the control description. "Reviewed by an independent agent" and "reviewed by a second model that reads the same transcript" are different controls with different residual risk. Your accountability record should say which one you have.
  • Push the residual to a mechanism that does not need to understand. If nothing can independently judge the action, bound the damage instead: a spend cap, a scope restriction, a reversible-only path. Blast radius is the control that does not depend on anybody being right.
  • Reserve the human for the correlated cases, not the frequent ones. A human reviewer is expensive and imperfect, but their c against an injected document is genuinely low, because they are not executing the instructions they read. Route by coupling, not by dollar value — the triage argument in human-in-the-loop.
  • Re-earn approvals, never replay them. Bind the verdict to the exact action — arguments hashed, a short expiry, single use. An approval that survives a retry, a resume or an argument change was never separated from the maker at all.

Do this today: take your highest-consequence agent action and fill in one row of the step-2 table honestly. If the context column is a ✗ — and it almost certainly is — try the cheapest fix first and give the checker only the resolved action, with no transcript, and see whether its disagreement rate moves. Then put the inequality from step 4 on the same dashboard as your approval rate, because the moment the two diverge you have a control that is still logging and no longer working. For the cases where that leaves you exposed, detecting agent compromise covers the signals that do not depend on a second opinion.