AI Blog

Cursor Projects vs Codex cloud vs Claude Code on the web vs Jules: buy the meter

Four cloud coding agents that look interchangeable on a feature table bill in four different shapes — a usage pool with overage, one allowance shared across every surface you use, a rate limit shared with the rest of your account, and hard task counts per tier — and each shape induces a specific, predictable misuse. Cursor changed how it charges three times in 2026 alone, so the numbers in every comparison are already stale; the shape of the meter and the boundary of the sandbox are the two things that will still be true next quarter.

By Agentic AI Wiki 13 min read

Pick a cloud coding agent on model quality and you have chosen the one thing you can change for free next month. The durable differences are the shape of the meter and the edge of the sandbox — and each meter shape induces a specific misuse: hard task caps make you over-stuff a task, a shared account allowance makes your background work starve your interactive session, and time-metered billing makes you pay most for the runs that are going worst.

At a glance

Five products, four meter shapes. Devin is in the table as the reference point for the fifth shape — billing by wall-clock of autonomous work — because it makes the others legible.

PlatformWhere the work runsMeter shapeThe misuse it induces
Cursor Projects Its own cloud machine, persistent across months Seat plus an included usage pool, overage billed in arrears Fan-out with no ceiling, discovered at invoice time
Codex cloud Isolated cloud environments, parallel tasks One credit allowance shared across web, CLI, IDE and chat Background batches eat the budget you code with
Claude Code on the web A cloud clone of the repo, one session per task Account-wide rate limit; no separate charge for the VM, and no separate budget Same as above, with the compute genuinely free — so nothing signals the trade
Google Jules A Google Cloud VM per task Hard task counts by tier, free at the margin within the cap Over-stuffed tasks, because a small task costs a whole slot
Devin (reference) Its own cloud workspace ACUs — roughly a quarter-hour of active autonomous work each You pay most for the runs that are stuck
What each meter is good and bad at Matrix with five rows — Cursor Projects, Codex cloud, Claude Code on the web, Jules and Devin — against four columns: fan-out is cheap, a stuck run is cheap, the budget is isolated from your other work, and the cap is predictable. Jules is strong on predictable caps and on a stuck run being cheap, but weak on fan-out. Claude Code on the web is strong on a stuck run being cheap and weak on budget isolation. Cursor Projects is weak on cap predictability because overage is billed in arrears. Devin is strong on budget isolation and weak on a stuck run being cheap. Meter shape, scored four ways Fan-out is cheap Stuck run is cheap Budget is isolated Cap is predictable Cursor Projects Weak Medium Medium Weak (arrears) Codex cloud Medium Medium Weak (shared) Medium Claude Code web Medium Strong (no VM fee) Weak (shared) Medium Google Jules Weak (slot each) Strong Strong Strong (hard count) Devin Medium Weak (time) Strong Medium Weak Medium Strong No row wins. Each column is a different thing the meter makes you responsible for.
No row is better. Each is a different thing to be careful about.

Four meters, four ways to hold it wrong

Cursor Projects — a pool with an overflow

Cursor's paid plans bundle a seat price with a dollar amount of included model usage, and bill the overage in arrears; Pro sits at $20 a month with $20 of included usage, Ultra at $200 with $400. Projects, shipped in beta on 10 September, puts a coordinator on top of that which delegates to as many subagents as the work needs.

That combination has no natural ceiling. A pool with arrears billing is not a cap — it is a speed bump with a bill attached — and the whole point of a coordinator is to increase the number of agents running. The control you need is one the product does not impose: a per-Project limit on in-flight delegated work, which is also the limit your review capacity implies. Treat the pool as a budget alert, never as a guardrail.

Codex cloud — one allowance, many surfaces

OpenAI moved Codex to credit-based pricing in 2026, and the allowance is shared: local messages and cloud chats draw from the same pool, with cloud work generally consuming more than the equivalent local turn, and a typical task landing somewhere in a single-digit-to-low-tens range of credits. Tasks can be started from the web, GitHub, GitLab, Linear or Slack.

The sharing is the trap. Kick off six background refactors on Monday morning and the budget they consume is the budget you were going to use to pair on the hard bug that afternoon. Nothing in the interface tells you that, because the two surfaces are presented as different products. The mitigation is scheduling, not spend: fire batch work when you are not working, which is also when its latency costs you least.

Claude Code on the web — free compute, shared budget

Web sessions clone the repo into Anthropic's cloud and run there; the same thing is reachable from the terminal, and with Git Live mode a session pushes to a branch and opens the PR itself. There is no separate charge for the VM — and, as the docs are explicit about, no separate budget either: rate limits are shared with all other Claude usage on the account, so parallel tasks consume proportionately more of it.

This is the cleanest meter of the four and the easiest to misread. Because the compute is genuinely free, the natural instinct is to run many sessions, and the only signal that you overdid it arrives as your interactive assistant slowing down or refusing. If you are going to fan out here, do it from a service account with its own limits rather than from the account you also work in — the same isolation argument that applies to any shared quota.

Google Jules — a hard count, and the stuffing problem

Jules has been generally available since 2025 and remains the most conventionally metered of the group: tiers from free to around $125 a month, with task counts rather than tokens as the unit. It clones into a Google Cloud VM, writes a plan, executes multi-file changes, runs tests and opens a PR; since early 2026 it can also read from a small curated set of MCP servers while planning.

A hard count is the only meter here that gives you real predictability, and it buys that with a specific distortion: since a one-line fix and a sprawling refactor cost the same slot, you are pushed to bundle. Bundling is exactly the behaviour that makes an unattended run fail — a bigger task has more ways to go wrong and a diff nobody wants to review. If you use Jules, the discipline is to keep tasks small anyway and treat the unused slots as the cost of a reviewable diff.

Devin — the meter that charges for confusion

Devin bills in ACUs, where one ACU is roughly fifteen minutes of active autonomous work — VM time, inference and bandwidth rolled into one number — charged only while the agent is actually working, at a couple of dollars each on pay-as-you-go. Small scoped tasks come in under an ACU; medium features run to several.

Metering wall-clock is the most honest reflection of what a cloud agent actually consumes, and it has the least comfortable incentive: a run that is thrashing costs more than a run that succeeds. That is the correct price signal and an unpleasant one, and it is why the governing control on a time-metered platform is a hard wall-clock ceiling per task, set low, with the agent instructed to return what it has.

The other durable axis: what the sandbox can see

What a cloud coding agent's sandbox can and cannot reach In the centre, a solid accent box is the cloud VM holding a clone of the repository at a chosen commit, a setup script and injected secrets. To its left, four reachable things: the repository tree, the dependency install, the test suite, and outbound network where policy allows. To its right, four unreachable things marked with dashed borders: local uncommitted files, a running dev server, local MCP servers, and anything on a private network. An arrow leaves the VM downward to a pull request. The boundary is the same everywhere Inside the boundary Repo tree at a commit Dependency install Test suite Outbound network, where policy allows Cloud VM Clone of the repository, your setup script, secrets you injected. Unreachable, everywhere Local uncommitted files A running dev server Local MCP servers Your private network A pull request the one integration point they all share Which half of your backlog is repository-shaped decides how useful any of these are — and that fraction is set by the boundary above, not by the model inside it.
The boundary is the same shape everywhere, and it decides which tasks are possible at all.

Every product here draws the same line, and OpenAI states it most plainly: a cloud task sees the Git host's copy of the repository and nothing else — no local files, no running dev server, no local MCP servers, no private network. Jules clones into a Google VM, Claude Code clones into Anthropic's, Cursor gives the Project its own machine. The wording differs; the reachable set does not.

That boundary, not the model, determines the task list. Anything verifiable by a test suite that installs from a public registry works everywhere. Anything that needs to hit an internal staging API, reproduce against a local database, or drive a dev server on your machine fails everywhere — and it fails late, after the agent has spent minutes discovering it. The practical split is not "which product is smarter" but which half of your backlog is repository-shaped, and that fraction is usually much smaller than the demo suggests.

Where the products genuinely differ is what you may put inside the boundary: a setup script, a container image, injected secrets, an egress policy. That is worth comparing carefully, and it is the same set of concerns as any other agent sandbox — see also egress control for agents, because a cloud VM with your repository and unrestricted outbound network is an exfiltration path whoever is running it.

What is actually expensive to change later

Three meter shapes and the distortion each one creates Three columns. Capped counts, as in Jules: spend is predictable, and the distortion is over-stuffed tasks because a small task costs a whole slot. Shared allowance, as in Codex cloud and Claude Code on the web: no separate compute bill, and the distortion is background work starving the interactive session. Metered work, as in Cursor's usage pool and Devin's ACUs: cost tracks real consumption, and the distortion is that a stuck or fanned-out run costs the most. Three shapes, three distortions Capped counts Jules Spend is predictable; a slot is a slot. Shared allowance Codex cloud · Claude Code web No separate compute bill and no separate budget. Metered work Cursor pool · Devin ACUs Cost tracks what the run actually consumed. The distortion The distortion The distortion You over-stuff tasks, and the diff gets unreviewable. Background batches starve the session you work in. The runs going worst, and the fan-outs, cost the most.
Numbers move every quarter. These three shapes have been stable for two years.

Comparisons of this category age badly for a reason that is worth stating: Cursor alone changed how it charges three times during 2026, which means most published plan tables describe a product that no longer exists. Anything you decide on the basis of a price per seat is a decision with a shelf life of about a quarter. Two things are not like that.

The first is the shape of the meter, because it determines the discipline your team has to maintain rather than the amount on the invoice. Capped counts require you to resist bundling. Shared allowances require you to separate identities. Metered work requires you to set wall-clock ceilings. Those habits outlive any specific price, and a team that adopts the wrong habit for its meter will be surprised in the same way every month.

The second is where the work integrates. All five open pull requests, which means your merge queue, your CI, and your reviewers are the shared downstream — and that is the part you cannot swap, because it is your own. A platform choice that doubles your PR arrival rate is a change to your review process wearing a procurement decision's clothes; the background-agent playbook is blunt that the review queue is the throughput limit, and it does not care which vendor is filling it.

When to pick which

SituationChooseBecause
Finance needs a number that does not move Jules A task count is the only cap here you cannot exceed by accident
You already live in one vendor's plan and want zero new procurement Codex cloud or Claude Code on the web Included in the subscription — at the price of sharing your own allowance
A long-running programme of mechanical, test-verifiable change Cursor Projects The coordinator is the only one that plans across months and fans out
You want cost to track the work honestly, per task Devin ACUs meter wall-clock, which is what a cloud agent really consumes
The work needs your private network or a running dev server None of them That task is outside the sandbox boundary everywhere — keep it local

If you are choosing for a team rather than for yourself, run the pilot against the meter rather than against the model: give one squad a fortnight, instrument PRs opened, PRs merged, reviewer-hours spent and material errors found after merge, and see which of the four distortions above shows up in your data. It will be one of them.

FAQ

Is a cloud coding agent better than the same model in a local CLI?

Different, not better. The cloud version buys parallelism and survives your laptop closing; it gives up your local files, your dev server, your private network and your local MCP servers. For repository-shaped, test-verifiable work the cloud wins on throughput; for anything that needs your actual environment the local CLI is not a fallback, it is the only option.

Which meter is cheapest?

The question does not survive contact with a real month, because the meters are not commensurable — a task count, a credit allowance and a quarter-hour of compute are three different units. What you can compare is cost per merged change that survived review, which you have to measure yourself and which usually reorders the candidates.

Does a shared account allowance really matter?

It matters the first busy afternoon. Background batches and your interactive session draw on the same budget, so the queue you started in the morning can degrade the assistant you are relying on at four o'clock. Running fan-out under a separate identity costs nothing and removes the interaction entirely.

Can I use several of these at once?

Yes, and teams do — the integration point is a pull request, which every one of them speaks. The thing that does not multiply is review capacity, so adding a second platform without adding reviewers just moves the queue, and the outcome shows up as a rising material-error rate on merged changes rather than as a visible failure.

Further reading

On this wiki:

Project sources: