Fail-Closed and Fail-Open

A36
Concepts · Agentic AI Explained

Fail-closed and fail-open.

Almost every control in a working agent stack fails open, and it does so in a way your dashboards read as a clean pass — the injection classifier that timed out returns nothing, and nothing looks exactly like "no injection found". That is the useful form of this concept: the interesting question is not whether a control should block or allow when it breaks, but whether the system can tell that it broke at all. Get that ordering right and the direction almost picks itself; get it wrong and you have a safety architecture whose outage is invisible.

STEP 1

Two directions, borrowed from a world where the control was a wire.

The terms come from physical safety engineering, where the failure mode was a property of the mechanism. A fail-closed door locks when power dies; a fail-open one releases so people can get out of a burning building. Notice what both assume: that the failure is detectable, because the thing either has power or does not.

  • Fail-closed (fail-secure). When the check cannot produce an answer, deny. You trade availability for containment: a broken authorizer stops work rather than waving it through.
  • Fail-open (fail-safe). When the check cannot produce an answer, allow. You trade containment for availability, on the theory that the thing being protected matters less than the thing being blocked.

Software inherited the vocabulary and quietly dropped the assumption. A network policy engine that returns an empty policy set is not obviously broken; it looks like a system with no applicable rules. An agent's tool-permission check that throws an exception swallowed three frames up is not obviously broken either — the call just proceeded. The direction is a one-line decision. Detectability is the engineering.

The distinction that keeps this concrete: a control's failure direction is only meaningful at a point where the control is the thing that decides. If the model is free to proceed regardless — because the "control" is a sentence in the prompt — then the control has no failure direction. It has a compliance rate. See guardrails for why that line is the one most teams have not drawn.

STEP 2

Agent stacks fail open by construction, three times over.

This is not a discipline problem. Three structural properties of the way agents are built push every default towards allow, and each one is worth recognising on sight.

  • An empty result means "nothing found", and "nothing found" means allow. This is the big one, and it is almost never coded deliberately. A redactor that matched no spans, a classifier that returned no labels, a retrieval filter that produced no denials, a selector list that matched no elements on the page — every one of these returns the same value on success-with-nothing-to-do as on failure. Any control whose interface is "return the list of problems" fails open the instant it errors, unless the caller distinguishes an empty list from an absent one.
  • Instructions have no error path. A rule that lives in the system prompt cannot fail, because it cannot run. It can only be outweighed — by a longer context, a more recent instruction, a plausible-sounding page the agent just read. From the outside this is indistinguishable from an enforcement control that is failing open on every request, which is why the instruction hierarchy is a preference rather than a permission check.
  • Agents retry, so fail-closed degrades into fail-open via the human. A loop that hits a denied check does not stop; it rephrases and tries again, burning budget, and eventually someone with production access turns the check off to unblock the release. A fail-closed control with no fast, legible recovery path has a half-life measured in incidents. This is the failure mode kill switches are designed around and the one retry policy keeps colliding with.
# fails open on every error, and the metric says 0 blocks — as designed
findings = injection_scan(doc)      # raises, caught somewhere above
if findings:                        # [] on a clean doc AND [] on a dead scanner
    block()

# the same control with a third state, which is the whole fix
verdict = injection_scan(doc)       # CLEAN | FINDINGS | UNAVAILABLE
if verdict.unavailable:
    degrade(reason="scanner down")  # a decision, recorded, with a metric

The second form does not say which direction to take. It makes the direction expressible, which is the prerequisite for taking either one on purpose.

STEP 3

Pick the direction from what the control is, not from how frightening the risk sounds.

"Fail closed on anything security-related" is the intuitive rule and it is wrong often enough to be dangerous, because it lumps together two kinds of component that behave nothing alike under failure.

  • Enforcement points fail closed, always. Authorization checks, tenant scoping in retrieval, the egress proxy, tool preconditions, spend and rate limits, the policy engine. These are deciders: there is no useful sense in which the action can proceed without them, and an unavailable authorizer that allows is not degraded — it is absent. Their availability is your availability, so give them a cache, a static fallback policy and a budget, rather than an exception handler.
  • Probabilistic detectors mostly fail open, and must say so. An injection classifier, a toxicity filter, an LLM judge, a PII scorer. These never had an authoritative answer to give; they shift a probability. Failing them closed converts a recall problem into an outage, and — because they sit on the hot path of every request — into the most self-inflicted denial of service available to you. Let them fail open and carry the fact forward: the request proceeds with the detector's verdict marked unavailable, which downstream may convert into a lower autonomy tier, a human review, or a refusal to take the irreversible step.
  • Irreversibility overrides both. The one place a detector's absence should stop the world is immediately before an action you cannot take back — a payment, a deletion, an email to a customer, a merge. Not because the detector is authoritative, but because the cost of being wrong stopped being symmetric. This is the same asymmetry designing for failure works with at the UX layer: closed on consequence, open on capability.

The sorting exercise takes an hour and it usually finds the same two bugs: an authorization check written as a soft filter in application code, and a detector on the hot path with no timeout, whose slow days are silently already fail-open because a five-second budget expired.

STEP 4

When you cannot fail closed, shrink what the open failure reaches.

The honest conclusion of Step 3 is that a large part of an agent's safety surface cannot be made fail-closed at acceptable cost. Detectors are not authoritative, the model's compliance is not a check, and every advisory control degrades to "allow" under load. That is not an argument for accepting the risk; it is an argument for making the open state survivable, which is a design move rather than a tuning one.

  • Bound the authority instead of the behaviour. If the injection scanner going down means the agent might act on a hostile page, the question is what acting on a hostile page can accomplish. Scoped, short-lived credentials and a default-deny egress list make the open failure boring — this is the entire argument of blast radius, and it is the only control in this list that keeps working when everything else is down.
  • Give every control a health signal separate from its verdict. Alert on "checks evaluated per thousand requests", not on "blocks". A detector whose block rate drops to zero is either a quiet week or a dead process, and the two look identical unless you count invocations. See evaluating guardrails and detectors.
  • Make the degraded mode a real product state. "Running without content checks — irreversible actions require approval" is a mode you can design, test and show a user. Silently continuing at full autonomy is what happens when nobody designed it, and it is the default. This is graceful degradation, applied to safety components rather than to model availability.
  • Test it by breaking things, not by reading code. Kill the policy engine, blackhole the classifier's endpoint, return malformed JSON from the scorer, and watch what the agent does. Nearly every fail-open surprise is found in thirty minutes this way and in none by review, because the code looks correct — the bug is in what an empty value means. Fault injection is the standing version of this.

Do the inventory in one sitting: list every check between a user's request and an effect on the world, and write two words against each — decider or signal, then closed or open. Fix the deciders first, since any decider that currently fails open is a live authorization bug rather than a resilience preference. Then take the top three signals and add a third return state plus an invocation counter, so their absence becomes something a dashboard can show and a downstream step can react to. Leave the direction of the signals alone until then: knowing that a control is off is worth more than either answer about what to do when it is, and it is the part you can ship this week.