Pricing latency.
If a person is waiting on your agent, forty seconds of p50 latency costs about $0.67 of a $60-an-hour employee, and your token bill for that task is probably four cents — so the cheaper model that takes twice as long is not a saving, it is a sixteenfold cost increase wearing a discount's clothes. Latency is a cost line, it belongs in the same table as inference spend, and once it is there the arguments about model selection resolve in the opposite direction from the one everybody assumes. The catch is that the cost is not linear in seconds, so you have to price the buckets rather than the average.
Two bills, one of which nobody adds up.
Every agent task generates an invoice you can read — tokens, tool calls, sandbox seconds — and a second cost that is real, attributable and never written down: the wall-clock time somebody spends waiting for it. Latency shows up on a reliability dashboard, gets discussed as user experience, and never appears in the business case. That is the whole failure, because a cost that is not in the table cannot lose an argument to one that is.
Written out for one blocking task, the arithmetic is embarrassing in how simple it is:
wait cost per task = wall-clock seconds the human is blocked / 3600 x fully loaded hourly cost of that human 40 s x $60/hr = $0.67 inference for the same task ~$0.04 ratio 16 : 1
- Use the fully loaded rate, not salary. Salary plus employer taxes, benefits, equipment and the allocated cost of the seat. For most knowledge roles that is roughly 1.25–1.4× base, and using base instead understates the whole argument.
- The ratio, not the figure, is the finding. Any blocking interactive agent with a sub-dollar token bill and a multi-second response is in this regime, and the ratio moves with the wage rather than with the model price — which means it has been getting worse every year that inference got cheaper.
- This is the missing half of unit economics. The usual model counts inference, platform and review. Wait time is the fourth line, and on interactive deployments it is frequently larger than the first three combined.
The cost is convex, so the average lies.
If wait cost were linear in seconds you could optimise the mean and stop. It is not, because the human on the other end does qualitatively different things at different durations, and the classic response-time thresholds — roughly a tenth of a second, one second, ten seconds — are thresholds in cost as much as in perception.
- Under ~1 second: nearly free. The person stays in the interaction. You are paying only the seconds themselves, and there is almost nothing to buy here.
- 1 to ~10 seconds: linear. Attention is held but unproductive. This is the region where the arithmetic above applies cleanly and where the seconds really are just seconds.
- Beyond ~10 seconds: a step, not a slope. The person context-switches away. You now pay the wait plus a re-entry cost — reloading the problem, re-reading what they asked, deciding whether the answer still applies — which is routinely minutes rather than seconds, and which does not shrink when the wait does. Crossing that line is the single most expensive thing your latency does.
- Beyond a few minutes: abandonment enters. Some fraction of tasks are never resumed. Multiply the whole task's value by the non-resumption rate and add it; on long-running work this term dominates everything else, and it is why async agent UX is an economic intervention rather than a cosmetic one.
Convexity is why the p95 is the number to budget against and the mean is the number to ignore. A p50 of 6 seconds with a p95 of 40 costs far more than a flat 12 seconds, because the tail is where the context switches happen — and the tail is not a rounding error on a workload that runs ten thousand times a day. Carry the distribution, the way forecasting agent spend does for cost, and get the distribution itself from measuring agent latency.
Four regimes, and the seconds are only worth money in two of them.
The mistake in the other direction is just as expensive: paying a premium for speed on a workload where nobody is waiting. Classify each deployment before you spend anything.
- Blocking interactive. A human waits, idle, for the answer. Price the seconds as in Step 1. This is the regime where a more expensive, faster model is almost always correct, and where it is almost never chosen.
- Async and batch. Nobody is waiting; the result lands in a queue, a report or an inbox. Seconds are worth approximately nothing and the correct optimisation is throughput and cost per token. Buying speed here is pure waste — and the move in the opposite direction, deliberately trading latency for a batch discount, is free money that most teams leave on the table.
- Realtime voice. Latency is not only a cost, it is a quality input: past roughly a second of dead air the caller assumes the line dropped. You are also billed per wall-clock minute on the telephony leg, so a slow turn costs twice. See the latency budget, which prices this in milliseconds rather than dollars.
- Deadline-bound. A nightly reconciliation that must finish before the business day. Latency has a cost of zero right up to the deadline and an enormous one immediately after. That is a capacity problem, not a speed problem — buy parallelism and headroom, per rate limits and provider capacity, and do not pay for a faster model.
Most organisations run all four and have one latency policy. That is the first thing worth fixing, and it costs nothing but a conversation.
What to buy, in order of return — and note where "faster model" lands.
Once seconds have a price you can rank interventions by dollars per second removed. The ordering that falls out is not the one teams start with.
- Delete steps. The only intervention that improves both bills at once and the only free one. A step that never runs takes no time and costs no tokens. Most agent loops contain at least one tool call whose result never changes the outcome — see agent cost control.
- Prompt caching. Also improves both bills, because a cached prefix is not reprocessed: time-to-first-token on a long stable prefix falls substantially alongside the price. This is the highest-return lever that requires no architectural change. See prompt caching.
- Parallelism. Independent tool calls issued together rather than in sequence, and independent sub-tasks fanned out. Costs the same tokens, removes wall-clock. See parallel tool calls.
- Streaming and progress. This one is peculiar and underrated: it does not reduce latency at all, it reduces the cost of latency, by keeping the person inside the sub-ten-second regime where they have not context-switched. Given the convexity in Step 2, it is frequently the largest dollar win available and it is a front-end change. See streaming and waiting and latency UX.
- A faster or larger model. Only now. It costs real money per call, and it is the intervention teams reach for first because it is the one that requires no engineering.
- Provisioned capacity and warm pools. Removes queueing and cold starts, carries a utilisation floor, and is worth it only when your own tail is dominated by waiting for capacity rather than by generation. See provisioned throughput and sandbox pools and cold starts.
Be honest about what speed costs, because three of the tricks move the bill sideways.
Latency reductions are not free, and several popular ones pay for themselves out of a budget that is kept in a different file.
- A smaller model is not reliably faster end to end. It is faster per token and frequently needs more steps, and a twenty-step task at 1.5 seconds a step beats an eight-step task at 4 seconds only if you do not count the extra failures. Measure time to completed task, never time to first response.
- Skipping verification buys seconds from the error bill. Dropping a checking step is the fastest latency win available and it moves money to the cost of being wrong, which is the larger account on most deployments. If you take this trade, take it explicitly and per task class.
- Reducing reasoning effort trades tokens for correctness. A thinking budget is a latency dial and a quality dial on the same shaft. Turning it down is legitimate; presenting it as a pure latency optimisation is not.
- Routing to a fast model raises variance. A router that sends 80% of traffic to a fast model and escalates the rest has a bimodal latency distribution, and the p95 the human experiences is the slow path plus the wasted first attempt. See model routing.
- Provisioned capacity converts a variable cost into a fixed one. Which is a real change to your cost structure and to your breakeven volume — see fixed costs and the pilot tax — not merely a discount.
One row on the dashboard, next to cost per successful task.
The version of this that changes decisions is a single figure per task class, computed from numbers you already have, sitting beside the token bill rather than in a reliability review nobody from finance attends.
fully loaded cost per completed task = inference + tools + platform # the invoice + (p50 blocked seconds / 3600) x loaded rate + P(wait > 10s) x re-entry cost + P(never resumed) x value of the task
- Per task class, never blended. The same agent serving an interactive surface and a nightly batch has two completely different numbers, and the blended one recommends the wrong model for both.
- Source the re-entry cost the way you source everything else here — from ten real cases. Ask five people what they do when the agent takes thirty seconds, and how long it takes them to pick the task back up. You will get a usable number in an afternoon and it will be defensible, which is more than the current implicit value of zero.
- Track the non-resumption rate as a first-class counter. Sessions abandoned mid-task, per class. It is the term with the widest spread and the one nobody instruments.
- Re-derive after every model change. A model swap moves both columns and almost always in opposite directions. The whole point of having the number is that the trade becomes arithmetic rather than taste — the same argument measuring ROI makes about the surrounding model.
Do the two-hour version this week. Take your highest-volume interactive task class, pull the p50 and p95 blocked seconds from tracing you already have, multiply by a loaded hourly rate your finance team will accept, and put the result in the same table as cost per successful task. In most interactive deployments the wait line is larger than the token line by an order of magnitude, and two decisions flip immediately: the faster model becomes obviously worth buying, and streaming the output becomes the highest-return engineering task on the board. Related: cost, quality & latency for the triangle, measuring agent latency for where the distribution comes from, and the cost of human review for the other place a person's hourly rate enters your agent's economics.