Fault injection: the outage your dashboard cannot see.
When a dependency fails inside an ordinary service you get a 500 and an alert; when it fails inside an agent you get a fluent, confident, wrong answer and a green dashboard, because the component handling the error is a language model whose whole training is to keep going. That is the failure mode fault injection exists to find, and it is why the discipline looks different here than it does in a microservice fleet: you inject at the tool boundary rather than the network, and you assert on the trajectory rather than the status code. The question is never "did it survive the fault" — it is "did it tell anyone".
The model is the thing that hides the fault.
A tool result reaches the model as text. A 200 with the right JSON, a 503 body, an empty result array, and a timeout stub the SDK inserted are all just tokens in the context window, and the loop's error-handling logic is a model that has been trained, hard, to produce a helpful next step. It does. That is the bug.
- Errors get absorbed, not raised. Search returns nothing, so the agent answers from parametric memory. The pricing tool times out, so the agent quotes a number from an example in its prompt. Neither path throws. Your error rate is unchanged and your answer is now wrong.
- Retries multiply the invisible cost. A transient fault the model works around three times is three times the tokens and three times the latency for the same task, with nothing in the metrics but a slightly fatter p99 — see timeouts and deadline budgets for how those constants compound.
- Classic chaos tooling aims at the wrong layer. Killing a pod or adding 200ms of jitter with
tcexercises your infrastructure's resilience, which is usually fine already. The interesting faults are semantic — a result that is empty, stale, truncated, or wrong-but-well-formed — and no network-level tool produces those. - The blast radius is downstream of the answer. An agent that improvises past a failed read then performs a write based on it has converted a dependency blip into a data-integrity incident. Read repairing agent side effects for what that costs to undo.
The one-sentence version: in a normal system, faults are loud and correctness is quiet. In an agent, faults are quiet and only correctness is loud — and you are not measuring correctness on every request. Fault injection is how you find out what the quiet path does.
The fault catalogue that actually matters.
Write this list down, because it is short, it is stable, and almost nobody tests past the first two entries. Every one of these happens in production weekly at any real volume.
- Empty success. The tool returns 200 and zero results. This is the single highest-yield injection because it is indistinguishable, in text, from "there is nothing to find" — and it is what a permission filter, an index lag, or a typo'd query all look like from inside the loop.
- Slow but successful. A call that returns correctly at 25 seconds instead of 400ms. It tests deadline propagation, not error handling, and it is the fault most likely to blow a run budget without tripping anything.
- Malformed and truncated output. JSON cut mid-object, a field that arrives as a string where the schema says integer, a 2MB response where you expected 4KB. Each of these takes a different path through your parser and a different path through the model.
- Stale success. Correct-looking data from before the last write. Nothing anywhere will flag this, which is why it belongs in a test suite rather than an alert.
- Error-in-a-200. The tool returns HTTP 200 with
{"error": "rate limited"}in the body. Your transport layer sees success; your model sees a sentence about rate limiting and decides what to do with it, unsupervised. - Provider-side refusal and content filtering. The model API returns a refusal or a truncated generation mid-run. Agents handle this astonishingly badly — they frequently re-ask, get refused again, and burn the loop budget.
- Schema drift. A field renamed, a new required parameter, an enum that gained a value. This is not hypothetical maintenance work; it is the steady-state behaviour described in third-party tool drift.
- Credential expiry mid-run. The token was valid at step 2 and is not at step 14. Long runs make this ordinary, and the correct behaviour — stop, surface, do not improvise a workaround — is rarely the observed one.
Inject at the tool boundary, deterministically.
The mechanism should be the least clever thing in your stack. A shim between the agent and its tools, controlled by a per-run configuration, that can rewrite or delay any tool result before it becomes text.
- Put it in the tool-invocation path, not the transport. A wrapper around your tool dispatcher — or a proxy your MCP client points at — sees the logical call with its arguments, which is what you need to target faults precisely ("fail the third
search_orderscall of this run, not a random 5% of HTTP requests"). - Seed everything. A fault plan is
(run_id, seed) → list of (call_index, fault). Given the same seed, the same run produces the same faults. A chaos test you cannot replay is an anecdote, and it will be dismissed as one — this is the same requirement as replay testing. - Carry it on the run, not on the environment. A header or run attribute that selects the fault plan means one shared environment serves both normal and injected traffic, and it means you can eventually turn this on for a small, labelled share of production.
- Inject at the model boundary too. Refusals, truncations, and 429s from the model provider need the same shim on the other side of the loop. Teams almost always build only the tool half and then are surprised by provider behaviour they never rehearsed.
- Never inject a write-path fault without idempotency. If a fault causes a partial write and you cannot reconcile it by key, you have created the incident you were testing for — the prerequisite is idempotency and retries, in place first.
fault_plan:
run_id: 7f3a… # replay handle
seed: 4291
faults:
- on: {tool: search_orders, call_index: 3}
inject: empty_success # 200, zero results
- on: {tool: get_pricing, call_index: 1}
inject: {delay_ms: 25000} # slow but correct
- on: {tool: post_refund, call_index: 1}
inject: {status: 200, body: '{"error":"rate limited"}'}
- on: {model: completion, step: 9}
inject: refusal
Assert on the trajectory, and name the three outcomes.
This is where agent fault injection diverges hardest from the microservice version. There is no status code to assert on, because the run succeeded — it produced an answer. What you grade is the path, and there are exactly three verdicts worth distinguishing.
- Correct-degraded. The agent noticed the fault, did something sensible — retried with backoff, used a fallback source, narrowed the task — and delivered a correct result or a partial one it described accurately. Ship this.
- Honest stop. The agent could not complete the task and said so, without inventing a result and without writing anything. This is a pass, and teams under-credit it badly; an agent that stops cleanly under a dependency failure is worth more than one that is clever three times out of four.
- Silent fabrication. The agent produced a confident, plausible, unsupported answer — or worse, wrote something. This is the only verdict that is a release blocker, and finding it is the entire point of the exercise.
The assertions that separate these are trajectory properties, not output strings: did any tool call retry, and with backoff? Did the final answer cite a source that the injected fault made unavailable? Did a write occur after a failed read in the same run? Did the agent's user-facing text contain an acknowledgement of missing data? Score them with the same machinery you use in trajectory evaluation, and file the failures into the buckets in failure taxonomy and triage.
One assertion pays for the whole programme: a write must never follow a faulted read in the same run. It is mechanical to check from a trace, it needs no judge, and it catches the class of bug that turns a five-minute dependency blip into a week of reconciliation.
Three venues, in order, with different budgets.
Fault injection is not one activity. Running it as one is how it ends up either too slow for CI or too dangerous for production, and then gets dropped.
- In CI, on a fixed suite. Ten to thirty scenarios, each pinning one fault to one call, asserted with deterministic trace checks only — no LLM judge, because the gate must be fast and stable. This catches regressions when someone changes a retry policy or a prompt's error-handling instruction. It belongs beside your other gates in eval-driven development.
- In staging, as a scheduled game day. Broader fault plans, multi-fault runs, longer horizons, and a human reading a sample of the trajectories. This is where you find the interesting behaviours — the agent that "solves" a permissions error by trying a different account, the one that treats an empty result as permission to widen its query until it matches everything.
- In production, narrowly and late. Only after the first two are boring. A capped share of traffic, internal or consenting tenants first, a live abort condition wired to your kill switch, and never on write-capable runs until the idempotency work is done. The reason to do it at all is that staging's tool fixtures are not your vendors, and the faults you have never seen are the ones the fixtures cannot produce.
Budget honestly: the CI suite should cost minutes, the game day a half-day of one engineer per month, the production programme a named owner. Compare that to load testing, which answers a different question — load tells you when the system falls over, injection tells you what it says while it is falling.
The scoreboard, and what improving it looks like.
Two numbers, tracked per release, both derived from the same runs:
- Silent-fabrication rate under injection. Of the injected runs, the share that ended in a confident unsupported answer or an unjustified write. This is the only number that has to trend to zero, and it is the one to put in front of a release review.
- Honest-stop share. Of the runs that could not succeed, the share that stopped and said so. Rising is good, and it is worth watching separately because the cheap way to drive fabrication down — making the agent refuse more — shows up here as an obvious trade you can then price.
When fabrication is too high, the fix is almost never a better prompt. It is structural: make the failure legible to the model with an explicit typed error rather than a stringified exception (tool error messages is the design work), and make the consequential action unavailable rather than discouraged — a tool the model cannot call under a degraded state beats an instruction telling it not to. Graceful degradation is the same argument from the availability side.
Start this week with one scenario and one assertion. Pick your highest-traffic tool, inject empty success on its first call in a copy of your ten most common tasks, and check one thing: does the final answer acknowledge that it found nothing? Most teams running this for the first time find at least one task where the agent answers confidently from memory, and that single result justifies the rest of the programme better than any argument on this page. Then add slow but successful and error-in-a-200, wire the three of them into CI with deterministic checks, and only after that go looking for a game day.