The Reference Monitor

A39
Concepts · Agentic AI Explained

The reference monitor.

Almost every control in a production agent stack fails at least one of the three properties that make a control trustworthy — and the one it usually fails is the one that matters, which is that the thing doing the checking must be somewhere the thing being checked cannot reach. That test is fifty years old, it takes about ten minutes to run against your own architecture, and it explains why a guardrail written into a system prompt is not a security control while the same rule expressed as a filesystem policy is. It also explains the price: every step you move a control further out of the agent's reach, the less it understands about what the agent was trying to do.

STEP 1

Three properties, and the question each one asks.

The reference monitor comes from James P. Anderson's 1972 Computer Security Technology Planning Study for the US Air Force, which set out what a mechanism must satisfy to be relied on for access control. The formulation has survived unchanged because it is not a design — it is a test, and it is cheap to run.

# Anderson's three requirements, as questions about your stack

complete mediation   Is there any path to the protected operation
                     that does not pass through this mechanism?
                     (if yes, the mechanism is advisory)

tamper-proof         Can the thing being mediated modify, disable,
                     bypass or lie to the mechanism?
                     (if yes, you have a suggestion, not a control)

verifiable           Is the mechanism small and simple enough that
                     you can convince yourself it is correct?
                     (if no, you cannot know whether it holds)

The three are not a wish list to be partially satisfied. They compose multiplicatively: a mechanism that mediates every call but sits inside the agent's own process is defeated by one successful code execution, and a tamper-proof mechanism with a second path around it is defeated by taking the second path. What makes the test worth running is that failures are usually obvious once you ask the question out loud, and almost nobody asks.

The third property is the one people quietly drop, and it has a specific agentic form: if your policy decision point is itself a language model, it is not verifiable in Anderson's sense — you cannot enumerate its behaviour, and it inherits every input-handling weakness of the system it is protecting. A model can be a very good detector. It cannot be the mechanism you are relying on, which is the argument developed in guardrails.

STEP 2

Score your own controls, and watch the prompt layer fail all three.

Take the controls you actually have and put them through the three questions. The exercise is short and the result is consistent across teams: the controls that feel most like safety are the ones with the worst scores, because they were written where it was easiest to write them rather than where they could be enforced.

# mediation / tamper / verifiable

"do not delete files" in the system prompt      ✗  ✗  ✗
an LLM judge screening the agent's proposal     ~  ✗  ✗
a check inside the tool-calling function         ✓  ✗  ~
an approval step the agent can skip on retry    ~  ✗  ✓
allow-list in the harness's own config           ✓  ~  ✓
seccomp / Landlock policy on the process         ✓  ✓  ✓
network policy at an egress proxy it cannot set  ✓  ✓  ✓
a watchdog on separate hardware                  ~  ✓  ✓

✓ satisfied   ~ partially   ✗ not satisfied

Two rows deserve attention. The in-prompt rule fails all three and is still the most common "control" in the industry: there is no path it necessarily intercepts, the text sits in a context the agent's inputs can influence, and it has no error path to test. The last row is the interesting one — a monitor on separate hardware is genuinely tamper-proof and genuinely verifiable, and it scores only a tilde on mediation, because it cannot see every operation. That is not a flaw in the product. It is the trade-off this whole subject is about.

Worth naming the second path explicitly, because it is where most real failures live: retries. A control that is checked once but the operation is retried, resumed or replayed through a different entry point has not mediated anything. Approvals are the classic case — see blast radius for what the unmediated path costs, and retry amplification for how the second attempt gets there.

STEP 3

Tamper-proof means a different trust domain — and each step outward costs vocabulary.

"Tamper-proof" is not a property of code quality. It is a statement about trust domains: the mechanism is tamper-proof with respect to a given adversary if that adversary cannot write to it. So the question is never "is this control well written" but "what would the agent have to compromise to reach it", and the answer is a ladder.

  • Same context. A rule in the prompt shares a channel with untrusted input. Compromise cost: one successful injection. This is instruction hierarchy territory, and the hierarchy is a preference, not a boundary.
  • Same process. A check in the tool wrapper. Compromise cost: any code execution inside the agent — which, for a coding agent, is the product.
  • Same host, different process. A policy daemon or a local proxy. Compromise cost: a local privilege escalation.
  • Same host, kernel-enforced. seccomp filters, an LSM policy, a network namespace installed before the agent starts. Compromise cost: a kernel bug. This is the first rung most teams can actually reach today.
  • Different machine. An egress proxy, a policy service, a broker in front of the tool. Compromise cost: reaching a system on a different trust boundary.
  • Different processor. A watchdog on a DPU or a separate management controller the host cannot address. Compromise cost: firmware.

Now the part that is usually left out. As you climb, the monitor's vocabulary gets coarser. A check inside the tool wrapper knows the user, the task, the argument values and the plan. A seccomp filter knows syscall numbers. An egress proxy knows destinations, methods and bytes. A DPU watchdog knows flows and DMA. The further out you push a control for tamper-resistance, the less of your intent it can express — which means the strongest available mechanism can only enforce the invariants you have managed to restate in its terms.

This is the reason "the model should refuse" and "the infrastructure should block" are not competing answers to the same question. They sit at opposite ends of a trade-off: the layer with the best understanding has the worst enforcement, and the layer with the best enforcement has the worst understanding. A real design uses both and is explicit about which invariants it is trusting to which — the mistake is believing you have one control when you have one of each, half-specified.

STEP 4

Restate the invariant until something can enforce it.

The practical method follows from the trade-off. For each thing you are actually relying on, find the outermost layer whose vocabulary can still distinguish the outcomes you care about, and put the control there. If no layer can express it, you do not have a control — you have an expectation, and it should be written down as one.

  • Translate intents into effects. "Do not exfiltrate customer data" has no reference monitor, because no mechanism outside the model can evaluate it. "This process may open TCP connections only to these three hostnames" does, and it is a large fraction of the same guarantee. The translation is the work; see egress control for agents and policy-as-code for agents.
  • Install the mechanism before the agent, not around it. A policy the harness applies at startup and cannot later relax satisfies tamper-proofness; the same policy the agent can reconfigure through a tool does not. This is what ambient authority looks like when you go looking for it.
  • Make the failure path the default. A mechanism that fails open has no mediation property at exactly the moment you need it, which is the whole subject of fail-closed and fail-open.
  • Log at the monitor, not at the caller. Records written by the mediated component are as trustworthy as that component. The decision log belongs to whatever made the decision — the reasoning behind decision receipts.
  • Keep the verifiable property honest. If your policy file has grown to a thousand lines of conditionals, you have traded away property three for coverage. Prefer a small deny-by-default core plus a short allow-list you can read in one sitting.

Do this today: write down the three invariants you would be most embarrassed to have violated, then for each one name the exact mechanism that would stop it and ask whether the agent could reach that mechanism. Most teams find that one is genuinely enforced, one lives in a prompt, and one is nobody's job. The prompt one is not necessarily wrong to keep — it is wrong to count. Then read sandbox and isolation patterns for the rungs you can reach without new hardware, and detecting agent compromise for what to do about the invariants that turn out to have no monitor at all.