Forty-nine percent of enterprise leaders have scaled back an AI agent deployment because costs outran value. Seven percent have an established ROI number. Put those two figures next to each other and the story stops being about agents being expensive: nine out of ten of the organisations that cut did it without a denominator. The measurement you cannot add afterwards is not the cost meter — it is the baseline you needed before the agent went in.
The four numbers
KPMG's Global AI Pulse for Q2 2026 surveyed 2,145 senior leaders across 20 countries at organisations above $50 million in annual revenue. Four findings from it have been quoted separately all week; they are much more interesting together.
| Finding | Share | What it actually measures |
|---|---|---|
| Scaled back, narrowed, delayed or paused an agent deployment after expected costs began to outweigh anticipated value | 49% | How many took an action. |
| Cite limited understanding of AI cost structures, including how token pricing works, as a deployment barrier | ~33% | How many know they cannot read the meter. |
| Have full, real-time visibility into what AI costs to run at scale | 26% | How many can read the meter. |
| Report established ROI | 7% | How many have the other half of the ratio. |
Note the phrasing of the first one: expected costs outweighing anticipated value. Both sides of that comparison are forecasts. This is not a finding that agents failed a cost-benefit test; it is a finding that half of large enterprises reached a point where the forecast stopped looking good and the only move available was to make the deployment smaller. And it happened in a quarter when the same research programme found multi-agent orchestration roughly doubling — from 9% to 18% of organisations in KPMG's US panel — and around three-quarters of leaders still naming AI a top investment priority. Nobody is walking away. They are flinching.
"Scaled back" is what you do when the only control is a switch
"Scaled back, narrowed, delayed or paused" is four words for the same underlying fact: the response was applied to the whole deployment. That is the coarsest control that exists, and it is the one you are left with when spend is attributable to an API key and a month rather than to a task class and an outcome.
The rungs above it are much better and much more demanding. Retiring a task class — noticing that document reconciliation costs $4.10 a run against a $1.80 manual baseline while ticket triage costs $0.11 against $6.00, and killing the first while expanding the second — requires that spend be joined to task type and to whether the task succeeded. Retuning a task class — the same run at a fifth of the cost through caching, a smaller model on the classification hop, and a step ceiling — requires all of that plus a trace you can look at. Neither is exotic. Both are simply unavailable to an organisation whose finest-grained cost fact is the invoice.
So the 49% is not evidence that these deployments were unprofitable. It is evidence that when the signal arrived, the only actuator wired up was the one that turns things off. A team that had cost-per-completed-task by workflow would have responded to the same signal by cutting 20% of the work and keeping 100% of the value, and would not have shown up in this survey at all.
Cost arrives on its own. Value has to be built.
Your model provider meters tokens because it needs to bill you. That meter is precise, timestamped, delivered monthly whether you asked for it or not, and denominated in a unit finance already understands. Nobody meters the value. There is no vendor whose revenue depends on telling you that the agent resolved 4,000 tickets that would have taken 900 human hours. If that number exists, someone at your company built it, on purpose, with effort.
The result is that every agent deployment produces a rising, legible, unbidden cost signal and a value signal that stays at zero until it is constructed. Give that arrangement two quarters and the emotional reality inside the business is that the agent is expensive and unproven — regardless of what the actual ratio is. The 7% ROI figure and the 49% pullback figure are the same phenomenon observed from two sides.
This is also why "we need better cost observability" is the wrong lesson to take from the survey, even though it is the one the survey most obviously suggests. Cost observability is the tractable half. The gateway logs are there, the usage APIs are there, the traces are there; attributing spend to a feature and a tenant is a week of work for a team that has decided to do it. The 74% who lack real-time cost visibility have a solvable problem. The 93% who lack an established ROI number have a harder one, because part of what they need was only available in the past.
The measurement you cannot take later
An ROI number is a comparison against a counterfactual: what this work cost, or how well it went, before the agent existed. Once the agent is handling the workflow, the counterfactual is gone. You can estimate it, and people do, and the estimate is exactly as credible as whoever is asking for the budget — which is why so few organisations claim an established number.
The baseline is cheap to capture and expensive to reconstruct, which makes it the highest-leverage thing on this list. Before an agent touches a workflow, record four things about the current process:
- Volume. How many of these happen a week, and how that seasonally varies. Without it, any per-unit figure you compute later is unanchored.
- Human time per unit. Measured, not recalled. The recalled number is systematically wrong, and it is wrong in the direction that flatters the project.
- Quality today. The error rate, the rework rate, the escalation rate of the human process. Agents are routinely held to a standard the incumbent process never met, and you cannot make that argument without the incumbent's number.
- The tail. What fraction of cases already take five times the median. That fraction usually predicts where the agent's cost will concentrate, and it is the first thing to carve out of scope.
Four numbers, a week of work, captured once. They are what turns "the agent costs $4.10 a run" from a scary fact into a comparison — and a comparison is the only thing that lets you cut a workload instead of a programme.
What to instrument before the next expansion
| Instrument | Answers | Effort |
|---|---|---|
| Pre-agent baseline: volume, human time, quality, tail | Is this workload worth automating at all? | A week, and only available before you start. |
| Cost per completed task, by task class | Which workload to cut. Failed runs are paid for twice and averages hide it. | Join usage to trace outcome. Days. |
| p95 and p99 cost per task, not the mean | Where the runaway loops are. The mean is set by the median run; the bill is set by the tail. | Free once the above exists. |
| Spend rate alert, not spend total | Whether something broke at 03:14, while you can still act. | Hours. |
| A per-tenant or per-workflow budget enforced in the request path | Nothing — it prevents rather than answers. It is what makes a bad week not a bad quarter. | Days, and it is the one that bounds catastrophe. |
The ordering matters more than the list. Teams reliably build the spend dashboard first because it is satisfying and visible, and then discover it tells them how much they spent without telling them on what, or whether it worked. Cost per completed task, split by task class, is the smallest instrument that supports a decision other than "less".
The case for the pullback
Some of that 49% is correct, and it is worth saying plainly. There genuinely are agent deployments whose economics do not work at any level of instrumentation — a workflow where a human takes ninety seconds and the agent needs a forty-step trajectory to reach the same answer is not going to be rescued by better dashboards. Pausing is the right call there, and "we tried it, it did not pay, we stopped" is a healthier outcome than the alternative where a project survives on narrative for two more years.
The asymmetry is in the reasoning, not the conclusion. A team that stops because cost per completed task exceeded the manual baseline on a workflow they measured has learned something durable and can apply it to the next candidate workflow. A team that stops because the monthly bill got uncomfortable has learned nothing, will re-litigate the same decision in six months when a vendor cuts prices, and has no way to tell the two cases apart. The survey cannot distinguish them either — which is worth remembering before treating 49% as a market verdict on agents rather than a snapshot of the instrumentation.
FAQ
Does this survey mean enterprise agent adoption is reversing?
No, and the same research says otherwise. Roughly three-quarters of leaders still name AI a top investment priority with spend holding steady, and in KPMG's US panel the share orchestrating multiple agents across workflows roughly doubled — from 9% to 18% — in a single quarter. The pullbacks are happening inside a market that is still expanding.
What is the single most useful metric to add?
Cost per successfully completed task, split by task class. Token price tells you nothing about whether the work got done, and average cost per run hides both the failures you paid for twice and the tail runs that dominate the invoice.
Is 49% a big number for this kind of survey?
It is large, but read it against the 7% who report established ROI. In a population where only one in fourteen has a defensible return figure, a near-majority reporting cost-driven pullbacks tells you mainly that the cost side of the ledger is the side that exists.
We already have a spend dashboard. Is that enough?
Only if it is joined to outcomes. A dashboard aggregated by API key and month reports at exactly the granularity you cannot act on. The question that drives a decision is "which workflow is unprofitable", and no amount of total-spend precision answers it.
We are already in production without a baseline. What now?
Reconstruct what you can from before the agent — ticket timestamps, throughput records, error and rework logs — and be explicit that it is an estimate. Then capture a proper baseline for the next workflow before you touch it. The second one is worth more than a heroic reconstruction of the first.
Is the answer just a cheaper model?
Usually not. A model at a fifth of the price that needs three times the steps re-sends a growing transcript on every one of them, and can cost more while being slower. Caching what repeats, capping what runs away and deleting steps that were never needed move the bill further than a model swap does.
Further reading
On this wiki:
- Scaling Back an Agent Deployment — how to cut a task class instead of a programme.
- Cost Attribution & Budgets — carrying the dimensions you decide on through every call.
- Measuring Agent ROI — the denominator, and why it has to be built.
- Agent Unit Economics — what a single run actually costs.
- Agent Cost Control — the four levers, in order of return.
- Cost & Quota UX — making the meter legible to the person who triggered the run.