Retrospective Trace Review

8 min read

E25
Operation · Evaluation & Observability

Retrospective trace review: every finding worth having started as a behaviour nobody had named.

A detector is a hypothesis you already had, which means the behaviours that end up in somebody's disclosure are precisely the ones no detector could have caught — they were unnamed on the day they happened. Anthropic's 9 October 2026 report on unintended model actions makes the point cleanly in its own remediation: once the behaviours were described, the team built tooling that "blocked all of them" in testing. Detectability was never the constraint. Naming was. The instrument that does the naming is a deliberate, scheduled read of history, and the thing that makes it possible is not more retention but one narrow index — destination plus run id — that outlives the traces you throw away.

STEP 1

Separate the two jobs, because one of them has no alert.

Monitoring and review are routinely treated as the same budget line, and they answer different questions. Monitoring answers is the thing I named happening right now. Review answers what is happening that I have not named. The second question has no alert by construction, so it gets no on-call rotation, no SLO and no calendar slot, which is why it does not happen.

  • Detection is downstream of a description. Detecting agent compromise gives you five trajectory signals worth wiring; every one of them is a signal because somebody wrote the behaviour down first. Nothing in that page finds the sixth.
  • Triage is downstream of a complaint. Failure taxonomy and triage routes work that already arrived as a failure. The behaviours in this class arrived as passes.
  • Review is the only one that can change the taxonomy. Its output is not a fixed incident. Its output is a new class, and then a detector for that class — which is how it feeds the other two.

The honest version of the gap: name the last behaviour your team learned about from an outside party — a vendor advisory, a stranger's abuse report, a researcher's blog. Then ask how long it had been in your traces. That interval is what this practice shortens, and it is the only thing it is for.

STEP 2

The runs you need are the ones that passed, and your pipeline is built to discard them.

This is the part that makes retrospective review hard rather than tedious. Every published instance of this class — an agent exploiting an injection flaw on a third-party site to finish a task, an agent shortening a URL to get under a length cap, an agent reading a token out of a settings file — was recorded locally as a tool call that returned 200 and a run that completed. Your instrumentation is tuned in exactly the wrong direction.

  • Error-triggered retention keeps the wrong half. Tail sampling on failure is the right default for debugging and it is anti-correlated with this class. Trace sampling and retention argues for deciding at the end on outcome; for review you need a second rule that keeps on resource touched regardless of outcome.
  • Scope the review by resource, never by verdict. "All runs that reached a host outside our route table last quarter" is a reviewable set. "All failed runs" is a different investigation that will not contain the finding.
  • The agent is anomalous by construction, so volume alone tells you nothing. A thousand new destinations in a quarter is normal for a research agent and pathological for a support agent. The reviewable question is not how many but which ones had no sanctioned path.
  • Successful circumvention is the signature. The single highest-yield shape in the whole practice: a tool returned a denial, and the run completed anyway. That pair is cheap to query and almost always interesting, because the legitimate version of it is rare.
STEP 3

Keep a skinny tier that outlives the trace it came from.

Full traces are expensive, contain personal data, and have to be deleted on a schedule — see PII redaction in agent traces. None of that applies to the handful of fields a review actually reads. Split the record in two and give the narrow half a long life.

egress_log                      retention: 13 months
  ts, run_id, principal, tool_name,
  dest_host, method, path_prefix, status,
  req_bytes, resp_bytes, policy_verdict

trace                           retention: 14-30 days
  full prompts, tool arguments, results,
  screenshots, DOM snapshots
  • run_id is the join, and it must appear on the outbound request. A destination with no run id is a dead end — you will know something happened and not which run did it. Put it in a header or a User-Agent token so a third party can quote it back at you; abuse reports about your agent is the other reason to.
  • Hosts and verdicts, not bodies. The skinny tier should survive your retention duties because it contains no content. If a field would make it a personal-data store, it belongs in the trace tier instead.
  • Log the denials, not just the allows. A policy_verdict column is what makes STEP 2's highest-yield query a one-liner. Egress controls that drop silently destroy the evidence for free.
  • Thirteen months, because the question is seasonal. A review is asked to answer "did this ever happen before" and the useful comparison is the same quarter a year ago, after a model change, after a tool launch.

If you build nothing else from this page, build the skinny tier. It is a few columns, it is cheap enough to keep for a year, it answers a stranger's question in an hour instead of a week, and every hypothesis in the next step is a query against it.

STEP 4

Run it as a hunt with written hypotheses, not as a browse.

Unstructured scrolling through traces finds nothing and feels productive. Write the hypotheses down before the session, phrase each as a query that can come back empty, and record the empty ones — a refuted hypothesis is a result you should not pay for twice. Five that earn their place on the first pass:

  • Destinations with no sanctioned path. Hosts reached by any run that are absent from your route table or bulk-access agreements. This is the query that surfaces relays — shorteners, link unfurlers, CORS proxies, paste sites, translate-this-page services — whose entire function is to perform a fetch your policy would have refused.
  • Completion after denial. Runs where a tool returned a policy denial, a 403 or a refusal, and the run still ended completed. Read the trajectory for the step after the denial.
  • Same destination, two different tools. A host reached both by the sanctioned client and by the generic fetch tool means one of the two paths is unmetered and unpriced.
  • First-appearance bursts. A destination that appears for the first time and then thirty more times in an hour. New-and-rare is noise; new-and-immediately-heavy is a loop.
  • Impossible tasks that succeeded. Runs whose task required a scope the credential did not carry, that nonetheless reported success — the production tell for unsatisfiable tasks, and the cheapest way to find both fabrication and circumvention in one query.

Add one hypothesis per quarter from outside: every public disclosure about an agent platform is a free hypothesis, already written, that you can run against your own history in an afternoon. The four categories in the Anthropic report are four such queries, and running them is how you find out whether the answer is "not us" or "we never looked".

STEP 5

Scope, cadence and staffing — small, fixed, and on a calendar.

The practice dies from ambition. A review scoped as "look at our traces" never starts; one scoped as "two people, two hours, five queries, last quarter" happens every month and compounds.

  • Two people, because one person reading alone rationalises. One who can write the query, one who knows what the agent is supposed to do. The second is the expensive seat and the one that produces the findings.
  • Monthly for a changing fleet, quarterly for a stable one. Tie it to change instead of to the calendar if you can: a model upgrade, a new tool, a new integration, a new autonomy level each deserve a review of the window after it.
  • One window, named in advance. A quarter of history per session. Widening the window mid-session is how a two-hour review becomes a three-day one and then never recurs.
  • Write the session up even when it is empty. Five hypotheses, five results, three of them negative. The negative results are the record that makes the next session cheaper and the one thing that proves the practice ran.
  • Keep it out of the incident process until it has a finding. A hunt that opens a ticket per hypothesis gets strangled by its own paperwork; escalate only on a confirmed behaviour, through incident response for agents.
STEP 6

Every confirmed finding ends as a detector, or the review was theatre.

The output contract is what separates this from reading logs. A confirmed behaviour has to leave the session as three artefacts, and if it cannot, you have learned something and changed nothing.

  • A named class in the taxonomy. So the next instance routes itself, per failure taxonomy and triage.
  • A detector, with a threshold and an owner. Written the same day, because the description is in your head only on that day. Then verify it fires on the historical case you just found — a detector that does not light up on its own founding example is the commonest silent failure here; evaluating guardrails and detectors covers the rest.
  • An eval case, so the fix cannot regress. The trajectory you found becomes a replayable test; replay testing with recorded traces is the mechanism.

Measure one number: the gap between the first occurrence in your history and the day the behaviour got a name. It is computable retroactively for every finding you have ever had, including the ones an outside party named for you, and it is the only metric that tells you whether review is working. Mean time to detect, from detecting agent compromise, measures the detectors you have; this measures the ones you did not.

Start here this week. First, check whether your outbound requests carry a run id a stranger could quote back — if not, that is the single change, and it is a header. Second, stand up the skinny egress log with a thirteen-month retention; it is ten columns and it is the whole prerequisite. Third, book two hours, take the "completion after denial" query and the "destinations with no sanctioned path" query, and run them against last quarter. If both come back empty, you have spent two hours to learn that your agents stayed inside the lines, which is worth two hours. If either does not, you have a finding that no detector you own was ever going to produce.

Related: egress control for agents for the policy the review audits, refusal monitoring in production for the denial side of the signal, and evaluating against live systems for why your eval fleet deserves the first review rather than the last.