Operations / Economics & ROI

Economics & ROI

The unit economics of agents — pricing, cost attribution, ROI measurement, build-vs-buy, failure modes.

  1. Build vs Buy vs Orchestrate
    Not a cost comparison but a question of which layer is your durable moat: the three-branch decision tree, the hidden costs each path omits, and the lock-in you price today but pay later.
  2. Agent Unit Economics
    Cost per token is the wrong unit; cost per successful task is the right one, with the success rate in the denominator where small reliability gains swing margin hardest.
  3. Per-customer economics
    Whole-system unit economics hide which customers cost you money — a per-tenant cost view, what drives the heavy-tail user, and the levers you actually have.
  4. Cost Attribution & Budgets
    The provider bill is at the wrong granularity to act on: tag spend by feature, tenant, user, and version, propagate it through fan-out, and make budgets runtime circuit breakers, not reports.
  5. Measuring Agent ROI
    Value over a defensible counterfactual, net of the human still in the loop, on a cumulative time-to-value curve — and why the "agent replaces a human" framing is a category error.
  6. Pricing & Packaging Agent Products
    Seat, usage, and outcome pricing each misalign somewhere; align price with delivered value but defend the floor, because more autonomy means you hold more variable-cost risk.
  7. Where the Economics Breaks
    Unit economics do not erode gradually — they invert at retry storms, the long tail, escalation, the eval bill, and the silent-failure tax; watch the failure surface, not the average.
  8. Provisioned Throughput & Commitments
    A 30% discount means breaking even at 70% sustained utilisation, which bursty agent traffic never reaches — so reserved capacity is a latency guarantee you can put in front of a customer, not a saving, and a term commitment quietly freezes your model choice in a market that moves quarterly.
  9. The Cost of Evaluation
    Eval spend scales with change rate, not traffic, and its price is set by the smallest regression you insist on catching — which costs quadratically, so a two-point threshold buys roughly 7,700 runs where a five-point one buys 1,400.
  10. The Cost of Human Review
    The reviewer costs ten to fifty times what the tokens do, and it is the only line that does not shrink when the agent gets better — because a reviewer has to read the correct outputs too; only calibrated selective review removes it.
  11. Scaling Back an Agent Deployment
    Nearly half of enterprise leaders cut an agent deployment last quarter over cost, and "scaled back, narrowed, delayed or paused" is four names for one blunt instrument — inside a single deployment one workflow can clear its manual baseline by 55× while another loses money every run, so cut the task class instead, carve out the tail before condemning the class, and write the resume condition down as a number.
  12. Free Tiers & Trial Economics
    A SaaS free tier is capped by human boredom; an agent free tier is capped by nothing, because the user is a loop and the marginal cost is real — so price the tier off the maximum a single account can consume rather than the average, meter work instead of days or seats, and enforce the ceiling at request time rather than in a monthly review.
  13. Forecasting Agent Spend
    An agent’s per-task cost is heavy-tailed, so the mean forecasts nothing and is biased low — worse at scale, because more volume draws more of the tail. Forecast a sum over task classes carrying each class’s p95, cap every class so the tail has a finite number, drive it off retry rate, context growth, route mix and cache hit rate, and reconcile each month into volume, mix and per-task drift.
  14. Price Deflation & Cost per Task
    Frontier input tokens cost about a sixth of GPT-4’s 2023 price and a cheap capable model a three-hundredth, yet almost nobody’s agent got six times cheaper, because agentic workflows burn 5–30× the tokens of a chat completion. The deflation lands on a unit you do not buy — which makes token-shaving a depreciating asset, a term commitment a directional bet against a three-year trend, and cost per successful task the only line worth putting on the dashboard.
  15. The Cost of Being Wrong
    You can read the token bill to four decimal places and nobody has ever computed the error bill, which on most deployments is one to two orders of magnitude larger — so price it as a product of error rate, escape rate, unit remediation cost and amplification, source the unit cost from the incidents you already had, and notice that detection latency is a multiplier you can buy down. The resulting figure is what sets autonomy per action and what finally prices a reviewer honestly.
  16. Fixed Costs & the Pilot Tax
    Agent economics are modelled as pure variable cost, so a pilot divides a total that is mostly standing bill — index, eval suite, trace retention, capacity floor, review roster, on-call — by a tiny task count and reports a number that says nothing about the agent. Name the six fixed lines, note that they move in steps triggered by audits and deprecations rather than by traffic, and report the breakeven volume instead of the cost per task.
  17. Pricing Latency
    Forty seconds of p50 on a task that blocks a $60-an-hour employee costs $0.67 against a four-cent token bill, so the cheaper model that runs twice as long is a sixteenfold cost increase wearing a discount’s clothes. The cost is convex — free under a second, linear to ten, a step change after that when the person context-switches away — so budget the p95, classify each deployment into one of four regimes before spending anything, and notice that streaming, which removes no latency at all, is often the largest dollar win on the list.
  18. Fractional Time Savings
    Forty people saving twenty minutes a day is thirteen FTEs on a spreadsheet and zero dollars in any budget anybody controls, because a fifth of a person is not a line item you can cut. Freed capacity becomes money only where it relieves the binding constraint, which gives exactly four auditable shapes — absorbed growth, cycle time on a priced clock, avoided external spend, a named deferred req. Meanwhile the cost side is fractional too and only one side gets instrumented, so delete the hours-times-rate line and see what survives.