Pick a cloud coding agent on model quality and you have chosen the one thing you can change for free next month. The durable differences are the shape of the meter and the edge of the sandbox — and each meter shape induces a specific misuse: hard task caps make you over-stuff a task, a shared account allowance makes your background work starve your interactive session, and time-metered billing makes you pay most for the runs that are going worst.
At a glance
Five products, four meter shapes. Devin is in the table as the reference point for the fifth shape — billing by wall-clock of autonomous work — because it makes the others legible.
| Platform | Where the work runs | Meter shape | The misuse it induces |
|---|---|---|---|
| Cursor Projects | Its own cloud machine, persistent across months | Seat plus an included usage pool, overage billed in arrears | Fan-out with no ceiling, discovered at invoice time |
| Codex cloud | Isolated cloud environments, parallel tasks | One credit allowance shared across web, CLI, IDE and chat | Background batches eat the budget you code with |
| Claude Code on the web | A cloud clone of the repo, one session per task | Account-wide rate limit; no separate charge for the VM, and no separate budget | Same as above, with the compute genuinely free — so nothing signals the trade |
| Google Jules | A Google Cloud VM per task | Hard task counts by tier, free at the margin within the cap | Over-stuffed tasks, because a small task costs a whole slot |
| Devin (reference) | Its own cloud workspace | ACUs — roughly a quarter-hour of active autonomous work each | You pay most for the runs that are stuck |
Four meters, four ways to hold it wrong
Cursor Projects — a pool with an overflow
Cursor's paid plans bundle a seat price with a dollar amount of included model usage, and bill the overage in arrears; Pro sits at $20 a month with $20 of included usage, Ultra at $200 with $400. Projects, shipped in beta on 10 September, puts a coordinator on top of that which delegates to as many subagents as the work needs.
That combination has no natural ceiling. A pool with arrears billing is not a cap — it is a speed bump with a bill attached — and the whole point of a coordinator is to increase the number of agents running. The control you need is one the product does not impose: a per-Project limit on in-flight delegated work, which is also the limit your review capacity implies. Treat the pool as a budget alert, never as a guardrail.
Codex cloud — one allowance, many surfaces
OpenAI moved Codex to credit-based pricing in 2026, and the allowance is shared: local messages and cloud chats draw from the same pool, with cloud work generally consuming more than the equivalent local turn, and a typical task landing somewhere in a single-digit-to-low-tens range of credits. Tasks can be started from the web, GitHub, GitLab, Linear or Slack.
The sharing is the trap. Kick off six background refactors on Monday morning and the budget they consume is the budget you were going to use to pair on the hard bug that afternoon. Nothing in the interface tells you that, because the two surfaces are presented as different products. The mitigation is scheduling, not spend: fire batch work when you are not working, which is also when its latency costs you least.
Claude Code on the web — free compute, shared budget
Web sessions clone the repo into Anthropic's cloud and run there; the same thing is reachable from the terminal, and with Git Live mode a session pushes to a branch and opens the PR itself. There is no separate charge for the VM — and, as the docs are explicit about, no separate budget either: rate limits are shared with all other Claude usage on the account, so parallel tasks consume proportionately more of it.
This is the cleanest meter of the four and the easiest to misread. Because the compute is genuinely free, the natural instinct is to run many sessions, and the only signal that you overdid it arrives as your interactive assistant slowing down or refusing. If you are going to fan out here, do it from a service account with its own limits rather than from the account you also work in — the same isolation argument that applies to any shared quota.
Google Jules — a hard count, and the stuffing problem
Jules has been generally available since 2025 and remains the most conventionally metered of the group: tiers from free to around $125 a month, with task counts rather than tokens as the unit. It clones into a Google Cloud VM, writes a plan, executes multi-file changes, runs tests and opens a PR; since early 2026 it can also read from a small curated set of MCP servers while planning.
A hard count is the only meter here that gives you real predictability, and it buys that with a specific distortion: since a one-line fix and a sprawling refactor cost the same slot, you are pushed to bundle. Bundling is exactly the behaviour that makes an unattended run fail — a bigger task has more ways to go wrong and a diff nobody wants to review. If you use Jules, the discipline is to keep tasks small anyway and treat the unused slots as the cost of a reviewable diff.
Devin — the meter that charges for confusion
Devin bills in ACUs, where one ACU is roughly fifteen minutes of active autonomous work — VM time, inference and bandwidth rolled into one number — charged only while the agent is actually working, at a couple of dollars each on pay-as-you-go. Small scoped tasks come in under an ACU; medium features run to several.
Metering wall-clock is the most honest reflection of what a cloud agent actually consumes, and it has the least comfortable incentive: a run that is thrashing costs more than a run that succeeds. That is the correct price signal and an unpleasant one, and it is why the governing control on a time-metered platform is a hard wall-clock ceiling per task, set low, with the agent instructed to return what it has.
The other durable axis: what the sandbox can see
Every product here draws the same line, and OpenAI states it most plainly: a cloud task sees the Git host's copy of the repository and nothing else — no local files, no running dev server, no local MCP servers, no private network. Jules clones into a Google VM, Claude Code clones into Anthropic's, Cursor gives the Project its own machine. The wording differs; the reachable set does not.
That boundary, not the model, determines the task list. Anything verifiable by a test suite that installs from a public registry works everywhere. Anything that needs to hit an internal staging API, reproduce against a local database, or drive a dev server on your machine fails everywhere — and it fails late, after the agent has spent minutes discovering it. The practical split is not "which product is smarter" but which half of your backlog is repository-shaped, and that fraction is usually much smaller than the demo suggests.
Where the products genuinely differ is what you may put inside the boundary: a setup script, a container image, injected secrets, an egress policy. That is worth comparing carefully, and it is the same set of concerns as any other agent sandbox — see also egress control for agents, because a cloud VM with your repository and unrestricted outbound network is an exfiltration path whoever is running it.
What is actually expensive to change later
Comparisons of this category age badly for a reason that is worth stating: Cursor alone changed how it charges three times during 2026, which means most published plan tables describe a product that no longer exists. Anything you decide on the basis of a price per seat is a decision with a shelf life of about a quarter. Two things are not like that.
The first is the shape of the meter, because it determines the discipline your team has to maintain rather than the amount on the invoice. Capped counts require you to resist bundling. Shared allowances require you to separate identities. Metered work requires you to set wall-clock ceilings. Those habits outlive any specific price, and a team that adopts the wrong habit for its meter will be surprised in the same way every month.
The second is where the work integrates. All five open pull requests, which means your merge queue, your CI, and your reviewers are the shared downstream — and that is the part you cannot swap, because it is your own. A platform choice that doubles your PR arrival rate is a change to your review process wearing a procurement decision's clothes; the background-agent playbook is blunt that the review queue is the throughput limit, and it does not care which vendor is filling it.
When to pick which
| Situation | Choose | Because |
|---|---|---|
| Finance needs a number that does not move | Jules | A task count is the only cap here you cannot exceed by accident |
| You already live in one vendor's plan and want zero new procurement | Codex cloud or Claude Code on the web | Included in the subscription — at the price of sharing your own allowance |
| A long-running programme of mechanical, test-verifiable change | Cursor Projects | The coordinator is the only one that plans across months and fans out |
| You want cost to track the work honestly, per task | Devin | ACUs meter wall-clock, which is what a cloud agent really consumes |
| The work needs your private network or a running dev server | None of them | That task is outside the sandbox boundary everywhere — keep it local |
If you are choosing for a team rather than for yourself, run the pilot against the meter rather than against the model: give one squad a fortnight, instrument PRs opened, PRs merged, reviewer-hours spent and material errors found after merge, and see which of the four distortions above shows up in your data. It will be one of them.
FAQ
Is a cloud coding agent better than the same model in a local CLI?
Different, not better. The cloud version buys parallelism and survives your laptop closing; it gives up your local files, your dev server, your private network and your local MCP servers. For repository-shaped, test-verifiable work the cloud wins on throughput; for anything that needs your actual environment the local CLI is not a fallback, it is the only option.
Which meter is cheapest?
The question does not survive contact with a real month, because the meters are not commensurable — a task count, a credit allowance and a quarter-hour of compute are three different units. What you can compare is cost per merged change that survived review, which you have to measure yourself and which usually reorders the candidates.
Does a shared account allowance really matter?
It matters the first busy afternoon. Background batches and your interactive session draw on the same budget, so the queue you started in the morning can degrade the assistant you are relying on at four o'clock. Running fan-out under a separate identity costs nothing and removes the interaction entirely.
Can I use several of these at once?
Yes, and teams do — the integration point is a pull request, which every one of them speaks. The thing that does not multiply is review capacity, so adding a second platform without adding reviewers just moves the queue, and the outcome shows up as a rising material-error rate on merged changes rather than as a visible failure.
Further reading
On this wiki:
- Background coding agents — designing runs nobody is watching.
- Sandboxing & execution — what belongs inside the boundary.
- Unit economics — cost per completed task, which is the only comparable number.
- The coordinator is the requester now — what changes when a trigger, not a person, starts the work.
- The local CLIs compared — the other half of this decision.