Fractional time savings.
Your agent saves forty people twenty minutes a day, which is thirteen full-time equivalents a year on a spreadsheet and zero dollars in any budget anybody controls — and the reason is not that the saving is fake, it is that a fifth of a person is not a line item you can cut. Whether those minutes become money depends on exactly one thing: whether the freed capacity relieves the constraint that was actually limiting output. Get that question right and a modest agent shows up in the P&L; get it wrong and a genuinely useful agent produces a renewal conversation you cannot win.
The standard calculation, and the exact line where it stops being money.
Nearly every agent business case is the same four lines, and the first three are usually right.
# The business case everyone writes tasks/month 12,000 minutes saved per task 18 hours saved per month 3,600 x loaded hourly rate $65 = "annual benefit" $2.8M <-- this line is not money # What the same run actually produced 40 people, ~20 min/day each, spread across 3 teams, 2 cost centres and 1 department that was never the bottleneck.
The multiplication is arithmetically fine. What breaks is the implicit claim in the units: an hour saved is priced at the rate you pay for an hour, which is only valid if the hour was going to be bought or sold. Saved hours that are neither released nor resold have a market value of zero and a real utility that is not zero — and conflating those two is the single most common reason agent programmes lose their budget in year two despite working.
This is not an argument against measuring hours. It is an argument that hours are an intermediate metric, and the business case has to carry the conversion step explicitly, the way the rest of measuring ROI insists.
Fractional savings are real utility and non-fungible currency.
Three properties of a fifth of a person explain almost everything about how these programmes are received.
- They are not summable across people. Twenty minutes from each of forty people cannot be assembled into one engineer who does a different job. Aggregation is a spreadsheet operation, not an organisational one, and no manager can act on the aggregate.
- They land on the wrong owner. The cost of the agent sits in one budget; the minutes land in several others. Nobody who benefits pays, and the person who pays cannot point at a reduced number. This is the ordinary cost attribution problem with the sign flipped — the benefits are as unattributed as the spend.
- They are re-absorbed by default. The freed twenty minutes go to the next thing in the queue, which is usually valuable and always invisible. Unless you decide in advance what the capacity is for, the honest description of the outcome is "the same work, slightly less rushed" — which is worth having and is not a number.
Say the true thing rather than the bankable-sounding thing. "Forty people report the task is less unpleasant and the queue is shorter by mid-afternoon" is defensible, survives audit, and renews. "$2.8M of annual benefit" invites a finance partner to ask which cost centre went down, and when the answer is none, the credibility loss is charged against the next proposal too. The failure mode is catalogued in economics failure modes; this is its most common single instance.
Four shapes that convert, and the evidence each one needs.
Fractional savings become money only by hitting a binding constraint. There are four shapes where that happens, each with a specific piece of evidence you must be able to produce. If your deployment does not fit one of them, it may still be worth doing — but write it up as a quality or retention argument, not a financial one.
# The four convertible shapes
absorbed growth volume up, headcount flat
evidence: tickets/month up X% with the same roster,
over >=2 quarters, no quality regression
cycle time a queue with revenue or cost on the clock
evidence: p50 AND p90 time-to-resolution before/after,
plus the rate that depends on it
(conversion, SLA penalty, days-sales-outstanding)
avoided external spend contractors, BPO, overtime
evidence: an invoice that got smaller, or a renewal
signed at lower volume
deferred requisition a named open role, not a vague plan
evidence: the req number and the date it was closed
or pushed, signed by the hiring manager
- Absorbed growth is the strongest and the slowest. It needs two quarters and a flat roster to be visible, and it is the one shape that survives a sceptical auditor, because the counterfactual is a headcount request that did not happen. Start the volume-per-head series before you deploy or you will never be able to show it.
- Cycle time only counts where the clock is priced. Faster resolution in a queue nobody was waiting on is not a saving. Faster resolution on a queue with an SLA penalty, a conversion rate or a cash-collection date attached is a direct one, which is why collections, claims and onboarding queues are where agent economics look best.
- Avoided external spend is the cleanest and the smallest. An invoice is unarguable. It also caps out fast, and the moment you claim it you have made the agent's uptime a procurement dependency — price that against build vs buy.
- Deferred requisitions need a number, not an intention. "We would have hired two more" is not evidence; req 4471 pushed from Q1 to Q3 with the hiring manager's name on it is. This is also the shape with the most political cost, and it should never be claimed without the affected team's involvement — see worker consultation and co-determination.
The cost side is fractional too, and only one side gets instrumented.
Here is the asymmetry that turns a positive programme into a neutral one without anybody noticing. The agent's benefits arrive as fractions of many people's days. So do its costs — and the fractional costs are almost never counted, because there is no dashboard for them and no owner motivated to build one.
# Fractional costs, in the order they get forgotten review minutes/item x items, at the reviewer's rate (the reviewer is usually senior) verification time spent checking output that was fine rework time spent undoing output that was not escalation the handoff that arrives with no context context switch the interruption tax on the reviewer ops toil prompt tweaks, tool drift, eval upkeep # Net, not gross benefit = minutes saved x volume cost = review + verification + rework + toil report the difference, or you are reporting half a ledger
Two of these deserve naming. Verification of correct output is pure loss and it scales with volume rather than with error rate, so it dominates at high throughput and low error — exactly the regime you were aiming for. And the reviewer is senior, so the minutes you add are more expensive than the minutes you saved; the full treatment is in cost of human review, and the error-side arithmetic in cost of agent errors.
The net can be negative while every individual claim in the business case is true. Eighteen minutes saved per task against four minutes of senior review per task is still strongly positive at a $65 doer rate and a $110 reviewer rate — but at two minutes saved and three minutes of review, which is what a marginal use case looks like, the programme is destroying value and reporting a benefit. Compute the difference per use case, not per programme; a portfolio average hides the ones you should switch off.
Instrument the constraint before you deploy, not the minutes after.
Everything above is unmeasurable retroactively, which is why this step has a deadline. The window to establish a baseline closes the day the agent goes live, and the metrics you need are not the ones an agent platform gives you by default.
- Name the binding constraint in one sentence, in writing, before launch. "Our claims team cannot process more than 900 files a week because adjudication review is the bottleneck." If you cannot write that sentence, you do not yet know which of the four shapes you are aiming at, and the honest move is to run the pilot as a learning exercise with no financial claim attached.
- Baseline the constraint's units, not the agent's. Files per week per adjudicator, p50 and p90 cycle time, queue depth at 4pm, overtime hours, contractor invoice. Tokens and tasks are the agent's units and they tell you nothing about conversion — the distinction that unit economics turns on.
- Instrument the review path from day one. Time-to-approve per item and approve-without-change rate are the two numbers that decide whether the net is positive, and both are cheap to capture at the review UI. They also double as quality signals — see production feedback signals.
- Hold the roster flat for two quarters if you are claiming absorbed growth. A team that grows during the pilot has destroyed the only clean evidence for the strongest shape. This is a staffing decision, so it has to be agreed before launch, not discovered at the review.
- Keep the agent's own cost line separate and current. Fractional benefit against a rising token bill is the other way this goes quietly negative; the controls are in agent cost control and the forecast in forecasting agent spend.
Write the case in the form finance can audit.
The final move is a formatting discipline, and it is the one that gets programmes renewed. Separate the claim you can audit from the claim you believe, put the auditable one first, and state the conversion mechanism rather than implying it.
# The two-part write-up BANKABLE shape absorbed growth mechanism claims volume +31% on a flat roster evidence files/adjudicator/week, Q1 vs Q3 net of 4.1 min senior review per file value 2 deferred reqs (4471, 4502), $310k annualised confidence medium; one quarter of flat roster remains NOT BANKABLE, STILL TRUE 40 adjudicators report ~18 min/day returned p90 cycle time down from 6.2 to 3.4 days overtime requests down, no roster change claimed # The sentence that must appear "The hours saved are not claimed as a financial benefit. The claimed benefit is the deferred requisitions, which the hours made possible."
- Put the mechanism in the claim. Finance rejects "hours x rate" and accepts "req 4471 pushed two quarters". The mechanism sentence is what converts the first into the second, and omitting it is what makes a true claim look like a padded one.
- Publish the non-bankable section anyway. It is why the users defend the tool at renewal, and it is the honest record of what the programme did. Deleting it to look rigorous costs you your only political constituency.
- Re-run the arithmetic per use case, quarterly. Marginal use cases drift negative as volume grows and review load grows with it. The decommission decision is a normal part of the portfolio — see scaling back an agent deployment.
Do this today: take your current agent business case and delete the "hours x rate" line entirely. Whatever survives is your real case. If nothing survives, you have not found the constraint yet — go and write the one-sentence bottleneck statement from STEP 5 with the team that owns the queue, then decide whether this deployment relieves it. Then read per-customer economics for the same discipline applied to revenue, and cost of human review for the line most likely to be eating your benefit.