Timeouts & deadline budgets: the number nobody chose.
Every timeout in your agent was picked by someone who could not see the others, and the product of those independent choices — not any of them individually — is the worst case you actually ship. A thirty-second tool ceiling inside a twenty-step loop with two automatic client retries is a run that can legitimately last most of an hour, and nobody wrote that number down or agreed to it. Smaller constants will not fix it. The fix is to stop configuring timeouts at call sites and start passing a deadline down the run, because a budget composes and a constant does not.
The arithmetic nobody did.
A request/response service has one hop and one timeout, so the timeout is the worst case. An agent loop has a step count you do not control, and the loop multiplies every number inside it.
Work a realistic configuration through. Twenty steps, each a model call plus a tool call. The tool client is set to 30 seconds. The model SDK is left at its default — the OpenAI Python library times out after ten minutes and retries certain errors twice on its own, before your code ever sees a failure. The honest ceiling on that run is not "thirty seconds" and not "ten minutes". It is twenty times the sum of both, times the retry multiplier, and it is measured in hours.
Nobody intends this, because the defaults are inherited from three libraries that disagree about what "too long" means:
- No timeout at all. Python's
requestssets none by default; its own quickstart warns that "failure to do so can cause your program to hang indefinitely". An agent whose tool layer is built on it has a step that can never end. - Five seconds.
httpxshipsTimeout(timeout=5.0)as its default configuration — reasonable for an API call, catastrophic for a tool that runs a test suite. - Ten minutes, plus two silent retries. The model SDKs, sized for a long reasoning request rather than for your loop.
Your agent inherits whichever of these fires first, which is rarely the one you would have chosen. And the tail is worse than the mean suggests, because the loop gives the slow path many chances to happen: a step that takes the slow branch one time in a hundred produces a slow run roughly one time in six at twenty steps. You do not get to average that away. The p99 of a run is set by the worst step in it, and you have twenty draws.
Write down the number you are actually shipping before you tune anything. Multiply your per-step ceiling by your step cap, add the model-call ceiling times the same cap, multiply by one plus your automatic retry count. If that product is larger than the SLO you published, you do not have a latency problem yet — you have a specification problem, and every fix below is downstream of fixing that.
A deadline is a value you pass. A timeout is a constant you configure.
Distributed systems settled this argument a decade ago and agent stacks have mostly not read the transcript. gRPC's own guidance puts it plainly: a deadline is "a point in time past which a client is unwilling to wait", a timeout is "the max duration of time that the call can take", and the two are related by adding the timeout to the clock at the moment the call starts. The deadline is the primitive. The timeout is a derived quantity.
Two mechanics are worth stealing verbatim:
- Propagation is automatic in the mature implementations. A server that calls another server honours the original client's deadline — enabled by default in Go and Java, explicitly in C++. gRPC's documentation is blunt about why you want this: doing it by hand is "error-prone".
- It travels as remaining time, not as a timestamp. gRPC converts the deadline back into a timeout with the elapsed time already deducted, precisely so that unsynchronised clocks on two machines cannot corrupt it. Your run's budget should move the same way.
Ported to an agent, that is four rules:
- Mint the deadline where you mint the run. The goal arrives, you allocate a run ID, and you allocate an absolute deadline in the same breath — they belong to the same record, for the same reason, and both need to reach every layer. If you are already carrying the run ID for trajectory reasons, the budget rides along for free.
- Derive every call's timeout as
min(remaining, ceiling_for_this_step_type). The per-step ceiling stays — it is how you stop one pathological tool from eating the whole run — but it can only ever shorten the call, never extend it past the deadline. - Refuse to start what cannot finish. If the remaining budget is below the p50 for this step type, do not make the call. Starting work you will abandon spends money, produces a side effect you will have to reconcile, and buys nothing. gRPC servers do this by cancelling an already-expired RPC rather than serving it; your harness should do it before dialling.
- Subagents inherit a slice, never a fresh allocation. This is the most common way the budget quietly evaporates: a parent with five minutes left spawns three subagents that each construct a default ten-minute budget, and the parent's deadline becomes decorative. A child's budget is a portion of the parent's remaining time, decided by the parent, and the parent must still be alive to receive the answer.
A timeout abandons work. It does not cancel it.
This is the part that turns a latency control into a correctness problem, and it is the one most agent stacks get wrong. gRPC's documentation says the quiet part directly: "the server application is responsible for stopping any activity it has spawned to service the RPC." Your timeout firing is a statement about your patience. It is not a statement about the other system's behaviour.
So when a tool call times out at thirty seconds and the remote handler commits at thirty-one, you are now holding a question rather than a failure: did it happen? And the agent's instinct — the model sees an error and proposes the call again — is the worst available answer, because the model has no way to know the first attempt landed. Two invoices, two tickets, two emails.
The controls are unglamorous and they all live below the model:
- Mint the idempotency key before the first attempt, not on retry. A key generated at retry time is a new key, which is the same bug wearing a hat. See idempotency & retries for the full contract.
- Make the timeout path a reconcile, not a retry. On expiry of a write, the next operation is a read by key — "did operation
kland?" — and only its answer decides whether to re-issue. Reconciliation is cheap, bounded and safe; blind retry is none of those. - Never let the model own the retry decision for a write. Reads are fine to re-drive from the loop. Writes are not, because the decision needs state the model does not have. The harness swallows the timeout, reconciles, and hands the model a settled fact. This is the same boundary tool error recovery draws between errors the model should see and errors it should never be shown.
- Record the abandonment. Every expired write is a potential orphan. Log it as one, with the key, so that repairing agent side effects is a query rather than an archaeology project.
Cancellation is a capability, not an assumption. If a tool can take minutes — a build, a migration, a long retrieval — it needs an explicit cancel endpoint and your harness needs to call it on expiry. Absent that, "timeout" means you have stopped watching a process that is still running, still costing, and still capable of changing your data. Durable execution engines exist largely to make this survivable; see durable execution.
Streaming and thinking destroy your liveness signal.
A read timeout asks "have bytes arrived recently?" For a decade that was a fair proxy for "is this making progress?" In an agent stack it is no longer either.
- Keepalives keep a dead run alive. An SSE stream that emits a comment frame every fifteen seconds will never trip a read timeout, no matter how thoroughly the thing behind it has stopped doing useful work. You have configured a liveness check on your transport and it is answering on behalf of a corpse.
- Reasoning tokens make time-to-first-useful-output variable by design. An adaptive effort budget is exactly a mechanism for spending unpredictable amounts of time before anything observable happens — see adaptive thinking & effort budgets. Any timeout tight enough to catch a hung request will also kill a legitimately hard one.
- Token arrival is not step progress. A model emitting a long, confident, wrong plan is producing bytes at full rate while the run goes nowhere.
Run two clocks instead of one, and make the first one semantic:
- An inactivity clock on state change, not on bytes. Reset it when something happens that a reader of the trace would call an event: a tool call emitted, a tool result returned, a message completed, a checkpoint written. Not on a keepalive, not on a token.
- A hard wall on the run deadline from STEP 2. Absolute, unconditional, and independent of how lively things look.
And give long-running tools a progress protocol — a heartbeat that carries a monotonically advancing unit of work, not just a pulse. Without one your only choices are killing work that was fine and waiting forever on work that is not, and you will pick wrong in both directions.
Reserve budget for the landing.
The most wasteful failure in a deadline-bounded agent is the run that spends 100% of its budget working and 0% of it answering. The deadline arrives mid-tool-call, the harness kills the run, and a nearly complete piece of work is returned as an error — having cost full price.
Treat the finish as a funded step. Reserve a tail — a flat twenty seconds, or 10% of the budget, whichever is larger — that nothing but termination may spend. Then spend the rest on a ladder rather than a cliff:
- At roughly 60% consumed, stop widening. No new parallel exploration, no new subagents, no speculative retrieval. Breadth is what you buy with budget you have; at this point you no longer have it.
- At roughly 80%, admit only calls that fit. Every remaining tool call must have a p95 below the remaining budget minus the reserve. This is the refuse-to-start rule from STEP 2, applied with a harder threshold.
- At the reserve, stop and land. Summarise what is established, what is not, and what was in flight when the budget ran out.
Two details decide whether this is useful or merely tidy. Tell the model the remaining budget. If it can adapt its effort, the budget is an input to that decision, not a wrapper around it — a model that knows it has ninety seconds left will pick a different plan than one that discovers it by being killed. And label the partial result as partial, with the boundary visible: "budget exhausted after 14 steps; the account reconciliation is complete, the vendor lookup was in progress and is not included." A truncated run presented as a finished answer is strictly worse than a clean failure, because it launders an unknown into a claim — which is precisely what a review queue cannot catch. Everything else in this step is ordinary graceful degradation; the labelling is the part people skip.
Operate it: four numbers, and one of them is not what you think.
Timeout policy is invisible on a standard dashboard, because everything it does shows up as errors and latency that look like somebody else's fault.
- Separate "we gave up" from "they failed". A client-side deadline exceeded and an upstream 503 are the same red bar and opposite problems: one is your configuration, one is their availability, and averaging them produces a quarter of work aimed at the wrong system. This distinction belongs in your failure taxonomy before it belongs on a chart.
- Attribute the exhaustion to the step that spent the budget, not the step that was holding it. This is the counter-intuitive one, and getting it wrong sends teams optimising innocent code. When a run dies at step 18, step 18 is rarely the culprit — the four-minute retrieval at step 3 is. Attribute by consumption share across the run, which requires per-step budget accounting you have to add deliberately.
- Watch the distribution of budget utilisation, not the count of timeouts. If p50 utilisation is 15% and p99 is 100%, your deadline is irrelevant to the median user and is the binding constraint for the tail — which tells you the fix is in the tail's cause, not in the number. Timeout count tells you none of that. Pair this with measuring agent latency, which is where the per-step distributions live.
- Alert on reconciliations, because that is the true cost of your policy. Every write you had to read back is a timeout that bought you uncertainty. If that rate rises, your ceilings are too tight, and tightening them further will make correctness worse while making the latency chart look better.
Then test it honestly. Load tests run with deadlines disabled measure a system you do not operate; the interesting behaviour is what the fleet does when budgets start expiring together, which is also when your concurrency assumptions and your provider's rate limits begin interacting badly.
Do this first, before touching a single constant: emit one field — budget remaining at every step boundary — into your existing traces, and leave it for a week. It costs an afternoon and it answers the three questions you are currently guessing at: whether your deadline binds at all, which step actually consumes the run, and how many of your "errors" are your own client giving up on work that finished. Almost every team that does this discovers their real ceiling is set by a library default nobody in the room knew about.
Related: kill switches for stopping a run that is inside its budget and still wrong, durable state & resumability for the runs that should be suspended rather than abandoned, and agent cost control for the other budget that the same step count multiplies.