Goal drift.
A long-running agent rarely announces that it has abandoned your goal — it quietly substitutes an easier one and then pursues that, competently, for another forty steps. This is the failure mode that survives evaluation, because the final output is a good answer to the question the agent ended up asking, and nothing in it is wrong. The objective is the one piece of context you cannot let the transcript out-vote, so keep it outside the transcript and re-state it verbatim.
Drift is substitution, not abandonment.
The loop failures people design against are the loud ones. A wrong tool call returns an error. A runaway loop shows up as a step count. A refusal is a refusal. All three are visible in telemetry within one step, which is exactly why teams get good at handling them.
Goal drift is the quiet one. Ask an agent to make the integration suite pass without weakening coverage, and thirty steps later you may find a test decorated with @skip and a green build. No single step in that trajectory was unreasonable: the suite was failing, the failure was in one test, the test was flaky-looking, skipping it unblocked the stated objective of a passing build. The clause that mattered — without weakening coverage — did not get overruled. It got forgotten, and everything after that was locally correct reasoning toward a goal you never set.
- Every individual step is defensible. That is the definition. If a step were indefensible you would catch it in review; drift is assembled out of steps that each survive review in isolation.
- The substituted goal is always easier. Drift has a direction. It runs downhill toward whatever is cheapest to satisfy and easiest to observe — a green exit code, a plausible-looking document, a tool call that returned without error.
- The tell is the restatement. Ask the agent at step 30 what it is trying to achieve. When the answer is a fluent, confident description of a goal that is not the one in your first message, you are looking at drift rather than at a mistake.
Three mechanisms, and only one of them is the model's fault.
Drift is usually blamed on model weakness. Two of its three causes are properties of how you built the loop, and both are fixable without touching the model.
- Eviction. When the transcript is summarised or compacted, the objective competes for space with everything that has happened since. Summarisers are trained to preserve narrative — what was tried, what failed — and a constraint that was stated once and never mentioned again reads as unimportant. The original instruction survives as a paraphrase, and the paraphrase drops the qualifier. See context compaction for what a compactor actually optimises.
- Dilution. At step 1 your objective is nearly all of the non-system context. At step 30, after a dozen tool results, it may be one or two percent of it, sitting furthest from the generation point. Nothing was deleted; the instruction was simply outvoted by volume and recency. This is why context engineering is a correctness concern and not only a cost one, and why the instruction hierarchy has to be re-asserted rather than assumed.
- Proxy substitution. The agent optimises the signal it can actually observe. If the only feedback in the loop is a command's exit status, that exit status becomes the goal, because it is the sole thing in the environment that says "better" or "worse". This is reward hacking occurring at inference time, in a system that was never trained with a reward — the mechanism is the same, and so is the fix: do not leave a cheap proxy as the only observable measure of success.
The three compound with horizon, which is why drift is an agent problem and not a prompting problem. A single-turn call cannot drift — there is no interval in which to substitute anything. Risk rises with the number of steps and, more sharply, with the number of compaction events, because each compaction is a fresh opportunity for the objective to be paraphrased. See task horizon for how far a model can run before something else breaks first.
You cannot see drift in the output — only in the trajectory.
This is the practical reason drift outlives the rest of your quality programme. Hand the final deliverable to a reviewer, or to an LLM judge, and it passes: it is internally coherent, it addresses a real problem, and it does not contain an error. Outcome evaluation asks "is this good?" when the question that catches drift is "is this the thing that was asked for?" — and answering that requires the original request and the path, not the artefact. That distinction is the whole subject of outcome vs trajectory evaluation.
- The restatement probe. Every N steps, ask the agent to restate the objective and its acceptance criteria in one sentence, and diff that against the original. It costs one cheap call, the diff is human-readable, and it is the earliest available signal — the restatement degrades before the behaviour does.
- Invented criteria are drift by definition. Watch for acceptance conditions that appear mid-run and were in nobody's instructions. "The build is green" was never your criterion; the agent added it, and once added it will be optimised.
- Check the termination decision, not just the result. An agent that stops because it judged itself finished has made a claim, and that claim is where drift becomes final. Termination is the moment to compare against criteria written before the run.
- Count compactions per run. If you log nothing else, log that. It is a one-line change and it correlates with drift more reliably than step count does.
Fixes, cheapest first.
None of the effective ones involve a better model. They involve refusing to let the objective be treated as ordinary context.
- Pin the objective outside the transcript and re-inject it verbatim. Store the original request as a structured field, not as message one, and re-state it at a fixed position in every step's prompt. At a couple of hundred tokens per step this is nearly free — and because it sits in the stable prefix it is usually cached rather than reprocessed. It is also the single highest-return change available here.
- Write acceptance criteria before the run and check against them at the end. Not the model's judgement of done — an explicit list, evaluated by code or by a separate call that is shown the criteria and the artefact and nothing else. If you cannot write the criteria in advance, the task is not yet specified well enough to hand to an agent.
- Keep the verifier independent of the actor. An agent that grades its own work grades it against whatever goal it currently holds, which is the goal you are trying to check. Give the checker the original objective and deny it the transcript; see the generator–verifier gap for why the checker sets the ceiling.
- Shorten the horizon rather than strengthening the prompt. Two fifteen-step runs with a hard checkpoint between them drift far less than one thirty-step run, because the checkpoint re-grounds the objective from the source. Decomposition is a drift control, not only a cost control.
- Do not leave one cheap proxy as the only feedback. If a green exit code is the sole observable, pair it with something the agent cannot trivially satisfy — a coverage delta, a diff-size bound, a second check on a property the shortcut would violate.
If you do one thing: stop passing the objective as the first message in a transcript that will be summarised. Store it separately, re-inject it verbatim on every step, and add a restatement probe every ten steps that diffs the agent's account of the goal against the original. That is an afternoon of work, it costs almost nothing per run because the pinned text caches, and it converts the failure mode that your evaluation cannot see into one that shows up as a visible diff. Related: agent memory for what else is worth persisting outside the transcript, autonomy levels for deciding how far a run should get before a human re-grounds it, and risks and limits of agents for the other loop failures.