Anthropic published a cost curve on 22 September that argues against its own top-of-range setting: on FrontierCode, Claude Opus 5.5 at the default medium effort scores 54.6% for roughly $0.80 a task, and at max it scores 54.4% for roughly $6.19 — eight times the money for a difference too small to call a difference. The same dial on Terminal-Bench 4.0 is worth roughly eight points. That is the finding to take away: effort is not a quality knob with a right setting, it is a per-workload measurement, and the level in the headline is not the level you should deploy.
At a glance
Opus 5.5 is also the first Claude model whose effort parameter does not default to high, which matters more than the price cut.
| Model | Price in / out per MTok | Default effort | Context / max output |
|---|---|---|---|
| Claude Opus 5.5 (22 Sep 2026) | $4 / $20 | medium | 1M / 128K |
| Claude Opus 5 | $5 / $25 | high | 1M / 128K |
| Claude Fable 5.1 | $10 / $50 | high | 1M / 128K |
| Claude Sonnet 5 | $2 / $10 | high | 1M / 128K |
Four points on that curve, in the order a procurement conversation should see them: medium at 54.6% for about $0.80 per task, high at 54.0% for about $1.09, xhigh at 51.4% for about $2.25, and max at 54.4% for about $6.19. Anthropic's own framing of the same data is that at default effort the model beats GPT-6 Astra's best score for roughly a fifth of the cost per task — a claim about the default, not about the ceiling.
What actually shipped on 22 September
The price move is the least interesting part: $4 and $20 per million tokens against Opus 5's $5 and $25, with cache reads at $0.20 and a fast mode at $8 and $40. Anthropic's stronger claim is that typical workloads cost about 40% less to run than on Opus 5, which arithmetic on the list price cannot produce — the remaining gap is behavioural, from finishing the same job in fewer tokens and fewer turns.
Three API facts do more work than the price:
- The default is
medium. Every other Claude model that supports effort defaults tohigh. The documentation states the consequence plainly: a request that omitseffortruns one level lower than it did on Opus 5. Swapping a model string is therefore a behaviour change, not a version bump, and it is the kind of change that shows up as a quality regression with no diff to blame. - Effort is not a thinking budget. It applies to every output token — prose, thinking, and tool calls. Lower effort produces fewer and terser tool calls; higher effort produces more of them, with plans and summaries around them. In an agent loop that is a change in what the agent does, not only in what it spends, and the docs are explicit that it is "a behavioral signal, not a strict token budget".
- Thinking cannot be turned off. Adaptive thinking is always on; a request that sets
thinking: {"type": "disabled"}returns 400 at every effort level. There is no non-reasoning mode to fall back to, so effort is the only cost control at the model layer.
If you are migrating, set effort explicitly on the way in rather than inheriting a default you did not choose. Setting it to the model's own default is a no-op behaviourally and a large improvement in legibility: the value in the request is then the value in your repository, which is the property that makes the next default change harmless. See unpinned vendor defaults for the general version of this.
Two tenths of a point is not a difference
The tempting read of the FrontierCode sweep is that max is worse than medium. It is not: 54.4% and 54.6% on one benchmark run are the same measurement wearing different rounding. Which is precisely the point — the score difference is unresolvable at this sample size and the cost difference is a factor of eight, so the only defensible conclusion is that the top of the dial bought nothing you can detect and charged for it anyway.
The xhigh dip to 51.4% deserves the same discipline. It could be real — overthinking on structured work is a documented failure mode, and Anthropic's own per-model guidance for an earlier Opus warned that max "can lead to overthinking" on structured-output tasks. It could equally be one unlucky run. You cannot tell from a single number, which is the entire argument of eval variance and statistical power, and the reason a headline score is mostly flake until someone reports a confidence interval next to it.
So treat the curve as a shape rather than as four facts. The shape says: on code-change work, the payoff from thinking harder is flat somewhere around the default, and everything above it is priced as if it were not.
And on the next benchmark, the same dial is worth eight points
Terminal-Bench 4.0 is the counter-example that keeps this from being an argument for cheapness. Anthropic's headline number there is 66.4%, against 55.8% for Fable 5.1 and 52.3% for Opus 5 — and that headline is not a default-effort number. At default effort the same model lands near 57–58% for roughly $3 a task; the published curve keeps climbing as you spend, and third-party readings of the chart disagree about which level the 66.4% sits at. The direction is not in doubt even where the exact rung is: on long-running terminal work, more thinking buys several points; on code changes and review, it buys noise.
That asymmetry has a mechanical explanation worth holding on to. Effort raises the number of tool calls as well as their depth, and tool calls are where a terminal task makes progress — reading state, running a command, checking the result. On a patch-and-review task the model already has the context it needs in the diff, so extra exploration has nothing to find and extra deliberation has nothing to resolve. Effort pays where the bottleneck is information the model does not yet have, and not where the bottleneck is a judgement it will make the same way at any level.
Effort multiplies through the loop, which is why it is not a per-call knob
Per-call reasoning about effort produces the wrong answer because the parameter lands in three places at once. It changes tokens per turn, it changes how many tool calls a turn contains, and — because a better-reasoned turn can finish work that would otherwise need another lap — it changes how many turns the episode takes. The first two push the bill up. The third pushes it down, and it is the one that produces Anthropic's claim of a 40% saving on a 20% price cut.
This is also the reason a naive "keep effort low to save money" policy backfires on agentic workloads: a level that saves 30% per turn and adds one failed lap out of four has lost. The failure is invisible in a per-request dashboard, because every individual request got cheaper, and visible in exactly one place — the cost of the tasks that finished.
Cost per completed task is the only unit that survives
Put the two levels side by side with a success rate attached, and the comparison stops being about token price. Take a task family where the cheaper level completes 62% of attempts at $1.00 and the next level up completes 85% at $1.35, and assume a failed attempt is retried once before a human looks at it:
# cost per completed task = attempt cost / success rate (single-shot) medium : $1.00 / 0.62 = $1.61 per completed task high : $1.35 / 0.85 = $1.59 per completed task # the "expensive" level is cheaper # now price the failures you actually pay for review : 15 min of an engineer at $90/h = $22.50 per failed task medium : 0.38 failures x $22.50 = $8.55 + $1.61 = $10.16 high : 0.15 failures x $22.50 = $3.38 + $1.59 = $4.97
The token price difference is 35%; the decision-relevant difference is a factor of two, and it is dominated by a number that does not appear on any model comparison page. This is the unit agent cost control is built around, and the reason a procurement spreadsheet keyed on dollars per million tokens ranks models in an order that has little to do with what they cost you.
Two mechanics to get right while you measure. Changing the top-level effort value between requests invalidates the prompt cache, so "dynamic effort" inside one conversation is more expensive than it looks; on Opus 5.5, Opus 5 and the Fable 5.1 family a per-message effort change (beta header mid-conversation-output-config-2026-07-01) preserves the cached prefix, and everywhere else you should pick a level per workload and hold it. And set a large max_tokens at the higher levels: it is a hard ceiling on thinking plus response together, so a generous effort level with a tight output cap produces truncation rather than depth.
The sweep, as a half-day of work
Anthropic's own documentation now tells you to do this rather than carry settings over from an earlier model, which is an unusual thing for a vendor to put in a migration guide and worth taking literally.
- Pin everything else. One model snapshot, one prompt, one tool set, one harness version. An effort sweep with a prompt change in it measures nothing.
- Sweep all five levels on 30–50 of your own tasks, not 5. Between-task variance dominates between-run variance, so more tasks beats more repeats on the same budget. Record tokens, turns, wall-clock and outcome per attempt.
- Report dollars per completed task with an interval, and stop at the cheapest level whose interval overlaps the best one. That rule is what stops the ratchet: without it, every sweep ends with someone advocating
maxon a 0.2-point lead. - Re-run it on model changes, not on a calendar. The level that won on Opus 5 is not a recommendation for Opus 5.5 — different default, different calibration, different tool-call behaviour.
If you ship one change this week, make it this: set effort explicitly in every request, log the level next to tokens, turns and outcome, and start reporting cost per completed task on your top three workloads. Everything else in this post is a reason that number will surprise you — an eightfold spread in price across a dial whose payoff is flat on one task family and steep on the next, with a default that just moved under everyone's feet.
FAQ
Is max effort ever worth it?
Yes, on task families where the published curve is still rising and on one-off work where an engineer's hour dwarfs the token bill. What the FrontierCode sweep argues against is max as a default posture — eight times the cost of the default for a difference that a single run cannot resolve. Anthropic's own escalation ladder is effort first, then a different model: its guidance is to move to Fable 5.1 when evals on Opus 5.5 at higher effort still fall short.
Does the lower default mean Opus 5.5 is weaker out of the box than Opus 5?
It means it runs one rung lower out of the box. On the published comparisons the default-effort configuration still beats Opus 5 at max effort on the coding benchmarks, at a fraction of the cost per task — but a migration that assumes identical behaviour from an identical request is assuming something the docs contradict.
Why not just cap thinking tokens instead?
That control is gone. Adaptive thinking is always on for this generation and thinking: {"type": "disabled"} returns 400 at every level; the older budget_tokens shape is not accepted. The remaining levers are effort, max_tokens as a hard ceiling, and an advisory task budget for the whole loop.
How many runs do I need before the sweep means anything?
Enough tasks, more than enough runs. Agent scores have large between-task variance, so 30–50 representative tasks at one repeat each beats 5 tasks at ten repeats for the same spend — and report a paired comparison against your current level rather than two independent averages.
Does effort change anything besides cost and quality?
Latency and tool-call shape. Higher effort means more tokens per turn and more tool calls, which lengthens the wall-clock of every turn; lower effort means terser calls and less preamble, which can read as an agent that skips steps. If your product surfaces the agent's reasoning to a user, the dial is also a UX parameter.
Further reading
On this wiki:
- Adaptive Thinking & Effort Budgets — the API surfaces across vendors, and the override behaviour.
- Unpinned Vendor Defaults — why the default that just moved is a class of problem, not an incident.
- Eval Variance & Statistical Power — why two tenths of a point is not a result.
- Agent Cost Control — cost per completed task, and the budgets that hold it.
- Price Deflation & Cost per Task — why list-price cuts keep failing to show up in your bill.