Shadow Mode & Dark Launches

8 min read

E20
Operation · Evaluation & Observability

Shadow mode & dark launches.

A shadow run measures an agent that never has to live with its own mistakes — nothing it broke stays broken, no user reacts to it, and every step it takes starts from a world the incumbent kept clean. That makes the result an upper bound rather than a forecast, and the gap grows with the number of steps, which is precisely where you were hoping shadow mode would reassure you. Run it anyway, and read it as a filter for disasters rather than as a prediction of quality.

STEP 1

What shadow mode is actually buying you.

The pattern is simple: real production input goes to the incumbent — a human, a rules engine, an older agent — and a copy goes to the candidate, whose outputs are recorded and discarded. You get the candidate's behaviour on the real input distribution, which is the thing no eval set ever has, with a blast radius of zero.

That distribution is the whole value. An eval suite is built from failures you already knew about, and it drifts toward pathological inputs; production traffic contains the boring middle, the malformed requests, the customer who pastes an entire email thread, and the seasonal shape nobody wrote a case for. A candidate that looks fine on 300 curated cases and falls over on 4% of real traffic is a completely ordinary outcome, and shadow is the cheapest way to find out before anyone is affected.

What it is not is a dress rehearsal. Three things are missing from a shadow run and all three are load-bearing: the agent's own errors never enter its future inputs, no human ever reacts to its output, and nothing it does changes the state of the world it reads next. Every over-optimistic shadow result traces back to one of those three.

STEP 2

For an agent, shadow is not free and not side-effect-free.

"Shadow" borrowed its reputation from classifiers, where running a second model costs a forward pass and touches nothing. An agent is not a classifier. It calls tools, and calling a tool is an effect even when you throw the answer away:

  • Writes disguised as reads. A search that logs the query, a CRM lookup that stamps "last viewed", a support tool that marks a ticket as touched, a memory layer that writes what it retrieved. Any of these makes the shadow run visible to the production system — and a memory write is the nastiest, because it contaminates the incumbent's own future context.
  • Quota and rate limits. Doubling traffic against a third-party API is the classic way a shadow launch becomes an incident in a system nobody was testing. Budget the shadow path its own credentials and its own limits so that exhausting them degrades only the experiment.
  • Real money. Shadow costs full token price for output nobody reads, and for multi-step agents that is not a rounding error. Sample rather than mirroring everything; 5–10% of traffic usually answers the question, and trace sampling and retention is the same arithmetic applied one layer up.
  • Notification side channels. A tool that emails, pages, posts to a channel or opens a ticket must be stubbed, not trusted to be harmless. This is the one that produces an apology.

So the honest definition is not "the agent runs with no effects". It is: these specific effects are suppressed, these are permitted, and here is what the permitted ones cost. Write that list down — it is the difference between a shadow launch and an unannounced deployment. Idempotency, retries and side-effect safety covers the machinery for drawing the line inside the tool layer rather than around it.

STEP 3

The bias is compounding, and it grows with trajectory length.

Here is the part that makes shadow results systematically optimistic for agents rather than merely noisy.

On a single-step task, shadow is nearly honest: the candidate sees the same input as the incumbent and produces a comparable output. On a multi-step task, the two diverge at the first decision — and after that point the shadow agent is running in a world the incumbent has been keeping tidy. It reads records the incumbent already corrected. It queries state the incumbent already advanced. It never encounters the mess its own step three would have created, because step three never happened.

The arithmetic is unkind. If the candidate's steps are independent at 97% each, a ten-step task succeeds about 74% of the time — but only if the errors compound. In shadow they do not compound, because the world resets under the agent at every step, so the observed per-task success rate approaches the per-step rate. You are measuring capability and reporting it as reliability. Task horizon is the concept that keeps this distinction visible.

There are two designs, and they can claim different things:

  • Live parallel. The candidate runs alongside production on live state. Cheapest to build, most contaminated: everything above applies at full strength.
  • Replay from trace. The candidate runs against a recorded snapshot — pinned tool responses, frozen state, replayed inputs — so its own actions can change what it sees next. Costlier, and it is the only version that can make a claim about a trajectory. Staging environments for agents is where the fixtures come from.

Rule of thumb: use live parallel to compare single decisions, use replay to compare runs, and never quote a per-task number from a live-parallel shadow.

STEP 4

Scoring without ground truth: adjudicate the disagreements only.

Shadow produces two outputs per item and no label. The temptation is to score agreement with the incumbent and call it accuracy. Do not: agreement measures similarity to something that is itself wrong some of the time, and it is highest exactly where both parties are trivially right, which is most of your traffic.

The efficient design is to label only where they differ. Agreement rate tells you how much work there is; the disagreement set is the only place information lives. Sample it, have a human adjudicate blind — shown both outputs without knowing which is which — and record the outcome in three buckets: candidate better, incumbent better, both acceptable. That third bucket is usually larger than anyone expects and is the single most useful finding, because it says the two are different rather than one being worse.

Two cautions. Adjudication has its own reliability, and if two of your reviewers agree on only 72% of items, a promotion decision built on one reviewer's opinion is noise — the constraint in annotation and labeling ops. And if you use a model judge to scale this, remember the judge has not solved the problem, only moved it: calibrate it against the human-adjudicated subset before you trust a scaled number, per LLM-as-judge for agents.

The disagreement set is also the best source of new eval cases you will ever get: real inputs, a demonstrated behavioural difference, and a human verdict attached. Harvest it deliberately rather than letting it expire in a dashboard.

STEP 5

What shadow cannot see at all.

Four things are invisible to a shadow run by construction, and knowing which they are stops you from over-reading a clean result.

  • Human reaction. Nobody reads shadow output, so you learn nothing about whether people trust it, override it, escalate it, or quietly stop using the product. Trust is a deployment property, not a model property.
  • Feedback loops. Once live, the agent's outputs become inputs — to its own memory, to downstream systems, to the humans whose behaviour it shapes. A shadow run is the only period in the system's life when that loop is open.
  • Adversarial adaptation. Attackers respond to what is deployed. A shadow agent is not deployed, so nobody is probing it, which means a clean shadow injection record says nothing about week two of production.
  • The cost of being wrong. Shadow counts errors; it cannot weigh them. One irreversible mistake can outweigh a thousand small wins, and only the live system has stakes. The cost of being wrong is where that weighting gets built.
STEP 6

Graduate on written criteria, through a dark launch in between.

Shadow is not the last stop before live. The useful intermediate is a dark launch: the agent runs on real traffic with its effects suppressed at the last mile rather than at the tool layer — it composes the reply, files the draft ticket, prepares the refund — and a human either sends it or does not. Now you get the one thing shadow cannot give you, a human reacting to real output, while the blast radius is still a click.

Write the promotion criteria before the run starts, or they will be negotiated afterwards against a number somebody already likes. A workable set:

promote when, over N days of shadow traffic:
  disagreement_rate        < threshold agreed in advance
  candidate_worse_share    < threshold   (of adjudicated disagreements)
  catastrophic_count       == 0          (defined before the run, not after)
  p95_latency              within budget
  cost_per_completed_task  within budget

and then, over M days of dark launch:
  human_send_rate          > threshold
  human_edit_distance      trending down

Two operational notes. Keep shadow running after you go live — a permanent low-rate shadow of the next candidate is the cheapest regression detector you will ever own, and it shares its plumbing with detecting quality regressions. And decide in advance what happens to the shadow output you recorded: it contains real customer data, it was produced by a system nobody reviewed, and it is subject to the same retention and redaction rules as any other trace, per redacting PII from agent traces.

Do this in the next sprint. One: write the effect list — suppressed, permitted, priced — and stub the notification tools first. Two: mirror 5–10% of traffic, not all of it. Three: score nothing but the disagreements, blind, in three buckets. Four: write the promotion thresholds down before you look at a single number. If you take one thing from this page, take this: a live-parallel shadow run can tell you a candidate is unfit, and it cannot tell you a candidate is ready. Related: online experiments for agents for what comes after promotion, simulated users for closing the reaction loop offline, and repairing what the agent already did for when an effect you thought was suppressed was not.