Operations / Economics & ROI
Economics & ROI
The unit economics of agents — pricing, cost attribution, ROI measurement, build-vs-buy, failure modes.
- Build vs Buy vs OrchestrateNot a cost comparison but a question of which layer is your durable moat: the three-branch decision tree, the hidden costs each path omits, and the lock-in you price today but pay later.
- Agent Unit EconomicsCost per token is the wrong unit; cost per successful task is the right one, with the success rate in the denominator where small reliability gains swing margin hardest.
- Per-customer economicsWhole-system unit economics hide which customers cost you money — a per-tenant cost view, what drives the heavy-tail user, and the levers you actually have.
- Cost Attribution & BudgetsThe provider bill is at the wrong granularity to act on: tag spend by feature, tenant, user, and version, propagate it through fan-out, and make budgets runtime circuit breakers, not reports.
- Measuring Agent ROIValue over a defensible counterfactual, net of the human still in the loop, on a cumulative time-to-value curve — and why the "agent replaces a human" framing is a category error.
- Pricing & Packaging Agent ProductsSeat, usage, and outcome pricing each misalign somewhere; align price with delivered value but defend the floor, because more autonomy means you hold more variable-cost risk.
- Where the Economics BreaksUnit economics do not erode gradually — they invert at retry storms, the long tail, escalation, the eval bill, and the silent-failure tax; watch the failure surface, not the average.
- Provisioned Throughput & CommitmentsA 30% discount means breaking even at 70% sustained utilisation, which bursty agent traffic never reaches — so reserved capacity is a latency guarantee you can put in front of a customer, not a saving, and a term commitment quietly freezes your model choice in a market that moves quarterly.
- The Cost of EvaluationEval spend scales with change rate, not traffic, and its price is set by the smallest regression you insist on catching — which costs quadratically, so a two-point threshold buys roughly 7,700 runs where a five-point one buys 1,400.
- The Cost of Human ReviewThe reviewer costs ten to fifty times what the tokens do, and it is the only line that does not shrink when the agent gets better — because a reviewer has to read the correct outputs too; only calibrated selective review removes it.
- Scaling Back an Agent DeploymentNearly half of enterprise leaders cut an agent deployment last quarter over cost, and "scaled back, narrowed, delayed or paused" is four names for one blunt instrument — inside a single deployment one workflow can clear its manual baseline by 55× while another loses money every run, so cut the task class instead, carve out the tail before condemning the class, and write the resume condition down as a number.
- Free Tiers & Trial EconomicsA SaaS free tier is capped by human boredom; an agent free tier is capped by nothing, because the user is a loop and the marginal cost is real — so price the tier off the maximum a single account can consume rather than the average, meter work instead of days or seats, and enforce the ceiling at request time rather than in a monthly review.
- Forecasting Agent SpendAn agent’s per-task cost is heavy-tailed, so the mean forecasts nothing and is biased low — worse at scale, because more volume draws more of the tail. Forecast a sum over task classes carrying each class’s p95, cap every class so the tail has a finite number, drive it off retry rate, context growth, route mix and cache hit rate, and reconcile each month into volume, mix and per-task drift.
- Price Deflation & Cost per TaskFrontier input tokens cost about a sixth of GPT-4’s 2023 price and a cheap capable model a three-hundredth, yet almost nobody’s agent got six times cheaper, because agentic workflows burn 5–30× the tokens of a chat completion. The deflation lands on a unit you do not buy — which makes token-shaving a depreciating asset, a term commitment a directional bet against a three-year trend, and cost per successful task the only line worth putting on the dashboard.
- The Cost of Being WrongYou can read the token bill to four decimal places and nobody has ever computed the error bill, which on most deployments is one to two orders of magnitude larger — so price it as a product of error rate, escape rate, unit remediation cost and amplification, source the unit cost from the incidents you already had, and notice that detection latency is a multiplier you can buy down. The resulting figure is what sets autonomy per action and what finally prices a reviewer honestly.
- Fixed Costs & the Pilot TaxAgent economics are modelled as pure variable cost, so a pilot divides a total that is mostly standing bill — index, eval suite, trace retention, capacity floor, review roster, on-call — by a tiny task count and reports a number that says nothing about the agent. Name the six fixed lines, note that they move in steps triggered by audits and deprecations rather than by traffic, and report the breakeven volume instead of the cost per task.
- Pricing LatencyForty seconds of p50 on a task that blocks a $60-an-hour employee costs $0.67 against a four-cent token bill, so the cheaper model that runs twice as long is a sixteenfold cost increase wearing a discount’s clothes. The cost is convex — free under a second, linear to ten, a step change after that when the person context-switches away — so budget the p95, classify each deployment into one of four regimes before spending anything, and notice that streaming, which removes no latency at all, is often the largest dollar win on the list.
- Fractional Time SavingsForty people saving twenty minutes a day is thirteen FTEs on a spreadsheet and zero dollars in any budget anybody controls, because a fifth of a person is not a line item you can cut. Freed capacity becomes money only where it relieves the binding constraint, which gives exactly four auditable shapes — absorbed growth, cycle time on a priced clock, avoided external spend, a named deferred req. Meanwhile the cost side is fractional too and only one side gets instrumented, so delete the hours-times-rate line and see what survives.