Fixed costs, and why the pilot looks unaffordable.
Agents are sold as pure variable cost — pennies per task, scaling with usage — and that model is wrong in the direction that kills good pilots. A production agent carries a standing bill that does not move with traffic: the index, the eval suite, trace retention, the review roster, the on-call. At pilot volume that layer is most of the spend, so "cost per task" at 300 tasks a month is a measurement of your fixed costs wearing a variable-cost label, and teams cancel on it. Split the bill before you divide it.
The per-token model is a variable-cost model, and at low volume it is the smaller half.
Every vendor calculator, every internal business case, and most of the advice written about agent economics computes the same quantity: tokens consumed times price, plus perhaps a per-seat platform fee. That number is real, it is the one that scales linearly with usage, and on a pilot it is routinely a minority of what the deployment actually costs.
The reason is structural. Almost everything that makes an agent trustworthy rather than merely functional is provisioned once and paid for whether or not anyone uses it. You do not buy a half-sized eval suite because you are running a pilot. The vector index does not shrink when traffic is light. The reviewer who has to be available for the queue is available for an empty queue too.
- The test for which line an expense belongs on is one question: if traffic halves this month, does this bill halve? Inference spend does. Almost nothing else on the list does.
- Per-seat pricing is fixed, not variable, even though it is quoted per user. It moves with headcount on a quarterly cadence, not with usage, which means during a pilot it behaves exactly like a licence.
- The mistake is not optimism, it is category error. Teams are not wrong about the token price; they are dividing a total that contains both kinds of cost by a task count that only makes sense for one of them. See unit economics for the per-task frame this sits underneath.
Name the six standing lines, because unnamed they get attributed to the model.
The failure is rarely that these costs are unknown. It is that they are booked somewhere else — under platform, under data, under someone's headcount — and so the agent's cost per task looks either implausibly cheap or unexplainably expensive depending on which way the allocation fell.
- Retrieval and memory infrastructure. The vector store or search cluster, the embedding pipeline, and the storage under both. It is sized for corpus and latency, not for query volume, so it costs the same at 300 tasks a month as at 30,000. Re-embedding events sit here too — see reindexing and embedding migrations.
- The evaluation suite. Every CI run of the eval set is inference you pay for and no customer asked for, and its cost is set by the size of the set and the frequency of the runs. Cost of evaluation covers how it gets away from you; what matters here is that it is decoupled from production traffic entirely.
- Trace storage and retention. Agent traces are large, and a retention period set by compliance rather than by engineering turns them into a standing bill with a multi-year tail. Trace sampling and retention is the lever, and it is one of the few fixed lines you can genuinely tune.
- Provisioned or committed capacity. The moment you buy throughput to get latency guarantees or a rate-limit floor, you have converted a variable cost into a fixed one on purpose. That is often correct; it is never variable. See provisioned throughput and commitments.
- The minimum review roster. Human review is usually modelled per item, but staffing is not divisible: covering a queue during business hours has a floor of one person, and that floor is paid at any volume above zero.
- The engineering rota. An agent in production needs someone on call, someone to handle provider deprecations, and someone to re-run the evals after a model swap. This is the largest fixed line in most deployments and the one most often left out of the business case entirely.
Add the six honestly for a typical first production agent and the fixed layer lands somewhere between $8,000 and $30,000 a month before a single task runs — dominated by the last line, because one engineer is more expensive than everything else combined. Whether that is a lot depends entirely on volume, which is exactly the point: it is a number that says nothing about whether your agent is efficient.
Cost per task is a hyperbola, and pilots sit on the bad part of it.
Write it out and the shape of the problem is immediate:
cost per task = fixed / N + variable per task N = 300 fixed 15,000 -> 50.00 + 0.40 = 50.40 per task N = 5,000 fixed 15,000 -> 3.00 + 0.40 = 3.40 per task N = 50,000 fixed 18,000 -> 0.36 + 0.40 = 0.76 per task
- The same system, unchanged, spans two orders of magnitude in cost per task across that range. No model was swapped, no prompt was tuned, nothing got more efficient. Only N moved.
- At N = 300 you are measuring your fixed costs. The $50 figure is 99% amortisation. Comparing it to a human doing the task for $12 is not a comparison between an agent and a human; it is a comparison between a fixed cost divided by a tiny number and a marginal wage.
- The honest pilot metric is the breakeven volume, not the cost per task. "This becomes cheaper than the status quo above 4,100 tasks a month" is a decision-grade sentence. "It costs $50 a task" is not, and it has killed deployments that would have been profitable at production volume.
- The inverse error is equally common at scale. Teams running 200,000 tasks a month see fixed costs vanish into the third decimal place and stop tracking them — right up to the quarter they add a second region or a compliance retention tier and the "fixed" line doubles.
The fixed layer is a staircase, not a line.
Modelling fixed costs as a flat monthly number is better than ignoring them and still wrong, because they move in steps — and every step is triggered by something other than volume.
- A second on-call person. Triggered by coverage expectations, not by traffic. It is usually the largest single step any agent deployment takes, and it arrives when the first SLO is written rather than when usage grows. See SLOs and error budgets.
- A retention tier. Triggered by a compliance review or a customer contract. Going from 30-day to 7-year trace retention changes the storage bill by a factor that has nothing to do with how many tasks you ran.
- A second environment or region. Triggered by a data-residency requirement or a staging mandate. It duplicates most of the fixed layer at once — index, retention, capacity floor — while adding no capacity you needed. Staging environments is where this gets decided.
- A model migration. Triggered by a provider deprecation, on their calendar. It costs a full eval re-run, prompt work and a re-validation cycle, none of which is amortised over the tasks you happened to run that month. Model deprecation and migration is a recurring fixed cost dressed as an incident.
Forecast the steps by their triggers rather than smoothing them into a rate. A budget that says "fixed costs grow 4% a quarter" will be correct for two quarters and then wrong by 60% in the quarter the audit lands. This is the same discipline forecasting agent spend applies to the variable side's tail.
Build-versus-buy is a question about which kind of cost you want.
The split reframes the decision that teams usually argue on features. A managed platform is, economically, a machine for converting your fixed costs into someone else's, and billing you a margin for the service. That is a good trade or a bad one entirely depending on where you sit on the curve.
- Below breakeven, buying is almost always right — not because the platform is better, but because you are renting a fixed layer you cannot fill. Paying a 3× markup on a small variable bill beats paying 100% of a fixed one you use 4% of.
- Above it the arithmetic inverts, and the markup is now applied to a large number. This is the real content of build vs buy, and the crossing point is computable from the two lines rather than argued from preference.
- Beware the fixed costs a platform does not absorb. Buying a runtime does not buy you the eval suite, the review roster, or the engineer who owns the thing. Vendors quote against the variable line because that is the line they replace; the largest fixed cost stays with you either way.
- Scaling back does not return you along the same curve. Halving traffic halves the variable half and nothing else, so cost per task rises. Teams reducing an agent's scope are frequently shocked to find it got more expensive per unit — which is the mechanism behind several of the traps in scaling back an agent deployment.
Report it as two numbers and a threshold.
The reporting change is small and it is what makes the decisions come out right. One line becomes three.
- Fixed per month, variable per task, and the breakeven volume. Every agent deployment should carry these three figures, and a review should ask about the third before the second. The threshold is the number that answers "should this exist", and neither of the other two does.
- Show cost per task with the amortisation split out. "$3.40 per task: $3.00 amortised fixed, $0.40 marginal" prevents the entire class of error this page is about, and it costs one extra column.
- Track utilisation of the fixed layer explicitly. Tasks per month against the volume your current fixed layer could serve. Running at 6% of a provisioned commitment is a finding, and it is invisible in any per-task metric.
- Re-derive the split after every step change, not on a schedule. A new region, a retention change, a second reviewer — each one moves the breakeven volume, and a threshold nobody has recomputed since the pilot is worse than none.
- Do not let fixed costs hide inside "platform". A single opaque line item is how the eval suite, the index and the rota stop being managed at all. Break them out even when they are small; they are the ones that step.
Before the next pilot review, do this for one deployment: list the six standing lines with a monthly figure each, put the variable cost per task beside them, and solve for the volume at which the total beats the status quo. Then present the breakeven number instead of the cost per task. In most pilots the two numbers tell opposite stories, and the one that has been getting presented is the one that cannot support a decision. Related: measuring ROI for what the threshold feeds into, economics failure modes for the other ways the case goes wrong, and agent cost control for shrinking the variable half once the fixed half is honest.