AI Blog

The top of the dial bought nothing

Anthropic shipped Claude Opus 5.5 on 22 September with a cost curve that argues against its own ceiling: on FrontierCode the default medium effort scores 54.6% for about $0.80 a task and max scores 54.4% for about $6.19, while the same dial is worth eight points on Terminal-Bench. It is also the first Claude model that defaults to medium rather than high, so a model-string swap is a behaviour change. Effort is a per-workload measurement, and cost per completed task is the only unit that survives it.

By Agentic AI Wiki 13 min read

Anthropic published a cost curve on 22 September that argues against its own top-of-range setting: on FrontierCode, Claude Opus 5.5 at the default medium effort scores 54.6% for roughly $0.80 a task, and at max it scores 54.4% for roughly $6.19 — eight times the money for a difference too small to call a difference. The same dial on Terminal-Bench 4.0 is worth roughly eight points. That is the finding to take away: effort is not a quality knob with a right setting, it is a per-workload measurement, and the level in the headline is not the level you should deploy.

At a glance

Opus 5.5 is also the first Claude model whose effort parameter does not default to high, which matters more than the price cut.

ModelPrice in / out per MTokDefault effortContext / max output
Claude Opus 5.5 (22 Sep 2026)$4 / $20medium1M / 128K
Claude Opus 5$5 / $25high1M / 128K
Claude Fable 5.1$10 / $50high1M / 128K
Claude Sonnet 5$2 / $10high1M / 128K
FrontierCode score and cost per task across four effort levels Four effort levels on Claude Opus 5.5. Medium scores 54.6 percent at about $0.80 per task, high 54.0 percent at about $1.09, xhigh 51.4 percent at about $2.25 and max 54.4 percent at about $6.19. The score bars are nearly equal while the cost bars grow almost eightfold from medium to max. FrontierCode v1.1 — score against cost, Claude Opus 5.5 Score (% of tasks) Cost per task (USD) $2 $4 $6 medium default 54.6 $0.80 high 54.0 $1.09 xhigh 51.4 $2.25 max 54.4 $6.19 Score bars use a zero-suppressed axis from 40% to 60% so the spread is visible; cost bars start at zero. 8x the default, same score
FrontierCode v1.1, read off Anthropic's launch charts: the score is flat within noise and the bill is not.

Four points on that curve, in the order a procurement conversation should see them: medium at 54.6% for about $0.80 per task, high at 54.0% for about $1.09, xhigh at 51.4% for about $2.25, and max at 54.4% for about $6.19. Anthropic's own framing of the same data is that at default effort the model beats GPT-6 Astra's best score for roughly a fifth of the cost per task — a claim about the default, not about the ceiling.

What actually shipped on 22 September

The price move is the least interesting part: $4 and $20 per million tokens against Opus 5's $5 and $25, with cache reads at $0.20 and a fast mode at $8 and $40. Anthropic's stronger claim is that typical workloads cost about 40% less to run than on Opus 5, which arithmetic on the list price cannot produce — the remaining gap is behavioural, from finishing the same job in fewer tokens and fewer turns.

Three API facts do more work than the price:

  • The default is medium. Every other Claude model that supports effort defaults to high. The documentation states the consequence plainly: a request that omits effort runs one level lower than it did on Opus 5. Swapping a model string is therefore a behaviour change, not a version bump, and it is the kind of change that shows up as a quality regression with no diff to blame.
  • Effort is not a thinking budget. It applies to every output token — prose, thinking, and tool calls. Lower effort produces fewer and terser tool calls; higher effort produces more of them, with plans and summaries around them. In an agent loop that is a change in what the agent does, not only in what it spends, and the docs are explicit that it is "a behavioral signal, not a strict token budget".
  • Thinking cannot be turned off. Adaptive thinking is always on; a request that sets thinking: {"type": "disabled"} returns 400 at every effort level. There is no non-reasoning mode to fall back to, so effort is the only cost control at the model layer.

If you are migrating, set effort explicitly on the way in rather than inheriting a default you did not choose. Setting it to the model's own default is a no-op behaviourally and a large improvement in legibility: the value in the request is then the value in your repository, which is the property that makes the next default change harmless. See unpinned vendor defaults for the general version of this.

Two tenths of a point is not a difference

The tempting read of the FrontierCode sweep is that max is worse than medium. It is not: 54.4% and 54.6% on one benchmark run are the same measurement wearing different rounding. Which is precisely the point — the score difference is unresolvable at this sample size and the cost difference is a factor of eight, so the only defensible conclusion is that the top of the dial bought nothing you can detect and charged for it anyway.

The xhigh dip to 51.4% deserves the same discipline. It could be real — overthinking on structured work is a documented failure mode, and Anthropic's own per-model guidance for an earlier Opus warned that max "can lead to overthinking" on structured-output tasks. It could equally be one unlucky run. You cannot tell from a single number, which is the entire argument of eval variance and statistical power, and the reason a headline score is mostly flake until someone reports a confidence interval next to it.

So treat the curve as a shape rather than as four facts. The shape says: on code-change work, the payoff from thinking harder is flat somewhere around the default, and everything above it is priced as if it were not.

And on the next benchmark, the same dial is worth eight points

Two task families, two shapes of payoff from effort A comparison across three rows for two task families. On code-change work the score-versus-cost curve is flat past the default level, the bottleneck is judgement, and the advice is to stay at the default. On long-running terminal work the curve keeps rising, the bottleneck is information the model does not yet have, and the advice is to step up and measure. The same dial, two payoff shapes Code change and review (FrontierCode) Long-running terminal work Shape of the curve flat past the default still climbing What the bottleneck is judgement on known context information not yet gathered Where to start default, and justify any step up one level up, measured What to report dollars per completed task, with an interval, paired against your current level
Same model, same parameter, opposite advice — which is why the sweep has to run on your tasks.

Terminal-Bench 4.0 is the counter-example that keeps this from being an argument for cheapness. Anthropic's headline number there is 66.4%, against 55.8% for Fable 5.1 and 52.3% for Opus 5 — and that headline is not a default-effort number. At default effort the same model lands near 57–58% for roughly $3 a task; the published curve keeps climbing as you spend, and third-party readings of the chart disagree about which level the 66.4% sits at. The direction is not in doubt even where the exact rung is: on long-running terminal work, more thinking buys several points; on code changes and review, it buys noise.

That asymmetry has a mechanical explanation worth holding on to. Effort raises the number of tool calls as well as their depth, and tool calls are where a terminal task makes progress — reading state, running a command, checking the result. On a patch-and-review task the model already has the context it needs in the diff, so extra exploration has nothing to find and extra deliberation has nothing to resolve. Effort pays where the bottleneck is information the model does not yet have, and not where the bottleneck is a judgement it will make the same way at any level.

Effort multiplies through the loop, which is why it is not a per-call knob

Where the effort parameter lands in an agent loop One request-level effort setting affects three multipliers: tokens spent per turn, the number and verbosity of tool calls within a turn, and the number of turns an episode needs. The first two increase cost, the third can reduce it by finishing the task in fewer laps, so the net effect on cost per completed task is not determined by the parameter alone. output_config.effort one value per request 1. Tokens per turn thinking plus prose — cost up cost ↑ 2. Tool calls per turn more calls, more preamble — cost up cost ↑ 3. Turns per episode fewer laps to finish — cost down cost ↓ Net effect is a property of the task family, not of the parameter. Measure dollars per completed task, at the level you will deploy. A per-call view sees only the first two.
One parameter, three multipliers — and the third one can cut the bill instead of raising it.

Per-call reasoning about effort produces the wrong answer because the parameter lands in three places at once. It changes tokens per turn, it changes how many tool calls a turn contains, and — because a better-reasoned turn can finish work that would otherwise need another lap — it changes how many turns the episode takes. The first two push the bill up. The third pushes it down, and it is the one that produces Anthropic's claim of a 40% saving on a 20% price cut.

This is also the reason a naive "keep effort low to save money" policy backfires on agentic workloads: a level that saves 30% per turn and adds one failed lap out of four has lost. The failure is invisible in a per-request dashboard, because every individual request got cheaper, and visible in exactly one place — the cost of the tasks that finished.

Cost per completed task is the only unit that survives

Put the two levels side by side with a success rate attached, and the comparison stops being about token price. Take a task family where the cheaper level completes 62% of attempts at $1.00 and the next level up completes 85% at $1.35, and assume a failed attempt is retried once before a human looks at it:

# cost per completed task = attempt cost / success rate (single-shot)
medium : $1.00 / 0.62 = $1.61 per completed task
high   : $1.35 / 0.85 = $1.59 per completed task   # the "expensive" level is cheaper

# now price the failures you actually pay for
review : 15 min of an engineer at $90/h = $22.50 per failed task
medium : 0.38 failures x $22.50 = $8.55 + $1.61 = $10.16
high   : 0.15 failures x $22.50 = $3.38 + $1.59 = $4.97

The token price difference is 35%; the decision-relevant difference is a factor of two, and it is dominated by a number that does not appear on any model comparison page. This is the unit agent cost control is built around, and the reason a procurement spreadsheet keyed on dollars per million tokens ranks models in an order that has little to do with what they cost you.

Two mechanics to get right while you measure. Changing the top-level effort value between requests invalidates the prompt cache, so "dynamic effort" inside one conversation is more expensive than it looks; on Opus 5.5, Opus 5 and the Fable 5.1 family a per-message effort change (beta header mid-conversation-output-config-2026-07-01) preserves the cached prefix, and everywhere else you should pick a level per workload and hold it. And set a large max_tokens at the higher levels: it is a hard ceiling on thinking plus response together, so a generous effort level with a tight output cap produces truncation rather than depth.

The sweep, as a half-day of work

Anthropic's own documentation now tells you to do this rather than carry settings over from an earlier model, which is an unusual thing for a vendor to put in a migration guide and worth taking literally.

  • Pin everything else. One model snapshot, one prompt, one tool set, one harness version. An effort sweep with a prompt change in it measures nothing.
  • Sweep all five levels on 30–50 of your own tasks, not 5. Between-task variance dominates between-run variance, so more tasks beats more repeats on the same budget. Record tokens, turns, wall-clock and outcome per attempt.
  • Report dollars per completed task with an interval, and stop at the cheapest level whose interval overlaps the best one. That rule is what stops the ratchet: without it, every sweep ends with someone advocating max on a 0.2-point lead.
  • Re-run it on model changes, not on a calendar. The level that won on Opus 5 is not a recommendation for Opus 5.5 — different default, different calibration, different tool-call behaviour.

If you ship one change this week, make it this: set effort explicitly in every request, log the level next to tokens, turns and outcome, and start reporting cost per completed task on your top three workloads. Everything else in this post is a reason that number will surprise you — an eightfold spread in price across a dial whose payoff is flat on one task family and steep on the next, with a default that just moved under everyone's feet.

FAQ

Is max effort ever worth it?

Yes, on task families where the published curve is still rising and on one-off work where an engineer's hour dwarfs the token bill. What the FrontierCode sweep argues against is max as a default posture — eight times the cost of the default for a difference that a single run cannot resolve. Anthropic's own escalation ladder is effort first, then a different model: its guidance is to move to Fable 5.1 when evals on Opus 5.5 at higher effort still fall short.

Does the lower default mean Opus 5.5 is weaker out of the box than Opus 5?

It means it runs one rung lower out of the box. On the published comparisons the default-effort configuration still beats Opus 5 at max effort on the coding benchmarks, at a fraction of the cost per task — but a migration that assumes identical behaviour from an identical request is assuming something the docs contradict.

Why not just cap thinking tokens instead?

That control is gone. Adaptive thinking is always on for this generation and thinking: {"type": "disabled"} returns 400 at every level; the older budget_tokens shape is not accepted. The remaining levers are effort, max_tokens as a hard ceiling, and an advisory task budget for the whole loop.

How many runs do I need before the sweep means anything?

Enough tasks, more than enough runs. Agent scores have large between-task variance, so 30–50 representative tasks at one repeat each beats 5 tasks at ten repeats for the same spend — and report a paired comparison against your current level rather than two independent averages.

Does effort change anything besides cost and quality?

Latency and tool-call shape. Higher effort means more tokens per turn and more tool calls, which lengthens the wall-clock of every turn; lower effort means terser calls and less preamble, which can read as an agent that skips steps. If your product surfaces the agent's reasoning to a user, the dial is also a UX parameter.

Further reading

On this wiki:

Sources: