Trajectories

A32
Concepts · Agentic AI Explained

Trajectories.

An agent's real output is not the answer it hands back — it is the whole path it took to get there, and that path is the only place where a correct answer and a lucky one look different. Most teams never see it, because their tooling records model calls rather than runs, and a run has no row of its own. Fix the unit and half of your evaluation, debugging and cost problems turn out to have been the same problem.

STEP 1

A trajectory is the run, ordered — and the run is what you shipped.

In a chat system the unit of work is a pair: a prompt goes in, a completion comes out, and you can evaluate the pair on its own. In an agent loop that pair is a single step out of forty, and evaluating it tells you almost nothing. The unit that corresponds to "the thing the agent did" is the trajectory: the ordered sequence of everything between receiving a goal and terminating.

Concretely, a trajectory holds:

  • The goal it started from — the task string, the system prompt, the tool set it was given, the budgets it was given.
  • Every action and observation, in order — each tool call with its arguments, each result, each error, each retry.
  • The model's own intermediate output — the reasoning between steps where the vendor returns it, and what it said to the user along the way.
  • The terminal state — not just the final message, but why it stopped: goal reached, budget exhausted, error, human interrupt. These are four different outcomes and prose collapses them into one.

The final message is a summary of the trajectory written by the same system whose work you are trying to judge. That is a bad place to look for evidence, and it is where nearly everyone looks first.

Two words nearby that are not synonyms. A trace is the telemetry record of execution — spans, timings, parent-child links — and it is how you usually store a trajectory. A trajectory is the semantic object: goal, actions, observations, outcome. You can have excellent traces and still not be able to answer "what did run 4471 actually do", which is the symptom that the trace was designed for the infrastructure rather than for the agent.

STEP 2

The trajectory is where "right" and "right for the right reason" separate.

Two agents answer the same question identically. One retrieved the current policy document, quoted the relevant clause and checked the effective date. The other recalled something plausible from training, never called a tool, and happened to be correct. Score only the output and these are the same run. They are not the same system, and the difference shows up next week when the policy changes.

This is the practical reason trajectory evaluation exists, and it cuts in both directions:

  • Correct output, broken process. The lucky guess above. Also the agent that got the answer from an error message, the agent that succeeded because a retry masked a bug, and the agent whose subagent reported "nothing found" when its search tool was failing. All of these pass an output check and are unrepeatable.
  • Wrong output, sound process. Every step was the right step and the third tool returned stale data. Judged on output this looks like a model failure and someone will go and change the prompt. The trajectory says it is a data-freshness bug, and the fix is somewhere else entirely.
  • Neither, expensively. The agent reached the right answer at step nine and then kept going for thirty more steps because nothing told it to stop — the termination problem, invisible in the output and obvious in the path.

Note what this does not mean. Grading every step against an ideal path is usually a mistake: agents legitimately reach good outcomes by routes a human would not have picked, and a rubric that punishes that is measuring conformity. The useful questions about a trajectory are narrower — did it ground the claim it made, did it stay inside its budget, did it stop for the right reason, did it take an irreversible action it did not need to.

STEP 3

If nothing in your stack has a run ID, you do not have trajectories.

This is the part that is boring and decides everything. Most observability arrived from the LLM-API era, where the natural row is one model call, and a great deal of agent tooling inherited that shape. The result is a store full of steps with no reliable notion of the run they belong to — and a trajectory you cannot reconstruct is a trajectory you do not have.

The specific places the thread breaks, all of them common:

  • Retries and resumes. A run that crashed at step twelve and resumed lands as two records. Unless the ID survives the restart, your latency distribution now contains two short runs where there was one long one — and the ones that look fine are the ones that failed.
  • Delegation. A subagent is a separate model conversation and will be a separate trace by default. Without an explicit parent link, a failure inside a child is an unrelated run, and you have a credit assignment problem instead of a transcript.
  • Model swaps and fallbacks. The run that failed over to a second provider mid-flight is one trajectory, and most stacks will file it as two.
  • Everything that is not a model call. Cache hits, guardrail rejections, policy checks, approvals waiting on a human. These are events in the trajectory and they are the ones your incident review will want, precisely because no model call corresponds to them.

The fix is a single decision made early: mint a run ID when the goal arrives, attach it to every span, tool invocation, log line, spend record and child run, and make it the primary key of your agent data. Cost per task, time to completion, failure taxonomy and "show me what happened" are then all queries against one object instead of four reconstructions. This is what observability for agents means that is different from observability for services.

STEP 4

Trajectories are an asset and a liability, and both are larger than you expect.

Having decided to keep them, notice what you now hold.

They are the asset that makes everything else possible. A stored trajectory replays: you can run a new prompt against the same recorded observations, diff the paths, and see what changed without paying for the environment again. It is also the raw material for an eval set built from production rather than imagination, for a regression suite that reflects real user behaviour, and for the failure taxonomy that tells you which of your four problems is actually the expensive one. Teams that can answer "how often does this go wrong, and in what way" have trajectory data; teams that cannot, do not.

They are also the largest and most sensitive object your system produces. A trajectory contains everything the agent saw, which is by construction everything it was allowed to see: the customer record it read, the credential in the tool response, the document it retrieved. It is bigger than a chat log by orders of magnitude, it is subject to discovery, and it inherits none of your existing retention rules because nobody wrote them for it. The cost is real too — a full-fidelity trace of every run is a bill that arrives before the value does.

  • Redact at write time, in one place. A trajectory store is the most concentrated copy of your data that exists; see redaction in agent traces.
  • Sample the successes, keep the failures. Full retention on runs that errored, escalated or took an irreversible action; sampling on the rest. Nobody rereads a successful run.
  • Set a retention rule deliberately, not by default. Long enough to debug an incident and build an eval set; short enough that you are not holding a year of customer data you cannot justify.

Do one thing this week: take a run that went wrong in production and try to reconstruct it end to end from what you currently store — goal, every tool call and result, why it stopped. Time yourself. If it takes more than a few minutes, or you find yourself joining three systems by timestamp, the missing piece is a run ID, and adding one is a smaller change than any of the fixes you were considering instead.

Related: outcome vs trajectory evaluation for how to score one without punishing creativity, trajectory and process evaluation for the methods, and trace sampling and retention for what it costs to keep them.