Agent S3's README recommends UI-TARS-1.5-7B as its grounding model, which means two of the four most-starred open-source computer-use projects are not competitors — they are the top and bottom halves of one stack. Read the other two the same way and the comparison stops being a leaderboard: you are choosing which of four layers you are missing, and the only axis that genuinely separates them is whether a generative model decides where to click every single time the task runs.
At a glance
Four projects, four layers, and one of them is the only one that ships the part that decides where the pointer goes. Figures read from each repository on 9 October 2026.
| Project | Layer it owns | Licence | Ships the deciding part? |
|---|---|---|---|
| UI-TARS (ByteDance) | The grounding model | Apache-2.0 | Yes — weights, on Hugging Face |
| Agent S3 (Simular) | The policy around a model | Apache-2.0 | No — requires a main model and a grounding model |
| cua (trycua) | The machine, and fleets of them | Mostly MIT; see below | Partly — CUA-S1 small models for narrow decisions |
| OpenAdapt | A compiler from demonstration to program | MIT | Only at compile time — none at run time |
None of these four is a peer of another
A computer-use agent has three jobs that are usually described as one. Something decides what to do next and what to do when it fails: that is the policy, and it is mostly prompt engineering, planning and retry logic. Something decides where on a 1,440-pixel-wide screenshot the "Save" control is: that is grounding, and it is a separate model with a separate failure mode. And something owns the pixels, the filesystem and the clipboard: that is the machine.
Sort the four projects by that split and the comparison resolves itself. UI-TARS is a vision-language model trained end-to-end for the second job, published as Apache-2.0 weights; its repository reports 94.2 on ScreenSpot-V2 and 61.6 on ScreenSpot-Pro for the 1.5 generation, which are grounding benchmarks, not task benchmarks. Agent S3 is the first job, explicitly: its setup requires a main generation model and separately requires a grounding model, and the one it recommends is UI-TARS-1.5-7B. cua is the third job, at fleet scale, with a bench harness and trajectory export attached. OpenAdapt is none of the three — it is a fourth thing we will get to.
This is why the star counts mislead. UI-TARS Desktop carries 39.2k stars and cua 29.1k, against 12.6k for Agent S3 and 1.8k for OpenAdapt — but UI-TARS Desktop is an application shell around the model, and the model repository itself sits at 11.6k. Popularity here tracks how many people wanted to try a desktop agent, not how much of the problem each project solves.
Why the published scores cannot pick for you
Two independent things are moving inside the numbers in that chart, and both of them are larger than the gap between any two of these projects.
The first is which benchmark the word "OSWorld" refers to. OSWorld-Verified is near saturation — a paper citing it puts Claude Opus 4.8 at 83.5% and reads that as desktop computer use being largely solved. OSWorld 2.0, posted in July 2026 with 108 long-horizon tasks built around exactly the phenomena the earlier set underrepresented, scores the same model family at 20.6% binary completion with a 54.8% partial score at 500 steps, and has GPT-5.5 plateauing near 13%. A 63-point difference under one name is not a disagreement about agents; it is a benchmark generation boundary, and a score quoted without the version is unusable. See the benchmark landscape for how routinely this happens.
The second is the selection budget. Agent S3's headline is 72.6% on OSWorld at 100 steps, which its README notes surpasses human performance of roughly 72%. The same README reports 66% for the agent alone; the extra 6.6 points come from Behavior Best-of-N, which runs the task several times and picks. The pattern repeats across its other two benchmarks: WindowsAgentArena 50.2% rising to 56.6% when selecting from three rollouts, AndroidWorld 68.1% rising to 71.6%. That is an honest and clearly-labelled technique, and it is also a cost multiplier rather than a capability: three rollouts is three times the tokens, three times the wall clock, and three times the side effects unless every attempt is sandboxed. Compare it against a human baseline and you are comparing N tries to one — which is the whole subject of human baselines in agent evals.
Strip both effects out and what remains is too small to choose on. Pick on the layer instead.
The layer everybody needs and nobody wants to maintain
The machine is the unglamorous layer and the one that decides whether any of this reaches production, because an agent that drives a GUI needs a GUI, and a GUI that is also your laptop is not a deployment target. cua is the serious answer: VM-backed desktops on macOS, Linux and Omarchy, a driver that reaches native apps and browsers on macOS, Windows and Linux, Lume for local macOS and Linux VMs on Apple silicon, SDKs in four languages, and Cua Bench for building tasks and exporting trajectories for training. If you are evaluating or training rather than deploying one assistant, this is the project doing work nobody else here is doing.
Two caveats, both material. First, "MIT" is the repository's headline and not the whole licence: Cua Spaces — the desktop app, the part you would actually hand to users — is source-available under FSL-1.1-MIT and converts to MIT two years after each release, cua-som is AGPL-3.0-or-later, and the optional cua-perception extension is not MIT either. That is a reasonable structure, and it is also exactly the kind of detail that invalidates a licence-based procurement sign-off made from the badge. Second, neither cua nor any other project here publishes conventional GitHub releases, so pinning a version means pinning a commit or a component version string.
The cautionary case is the one most listicles still recommend. Bytebot — a self-hosted agent in a containerised Ubuntu 22.04 XFCE desktop with Firefox and VS Code, Apache-2.0, 11.1k stars, Helm and Compose deployment — was archived by its owner on 7 March 2026 and is read-only. It is a clean design and it is not a maintained dependency. If your shortlist came from a comparison post, check the archive banner before the feature table; this is the second category in a week where the most-linked project turned out to be frozen.
One of them takes the model out of the loop
OpenAdapt is the smallest project here by an order of magnitude and the only one that answers a different question. You demonstrate the task once; openadapt flow record captures it, openadapt flow compile turns the recording into a bundle, and the replay runs the compiled program without a generative-model API in the control loop. The compiler's job, in the project's own phrasing, is to mine the effect contract from the observed delta — it works out what the demonstration actually changed, and that becomes the thing the run is judged against.
Which produces the property that makes it worth your attention regardless of whether you adopt it: it reports VERIFIED only if an independent check agrees. The tutorial confirms a saved record through a read-only API that the writing screen never touches, and a run that cannot establish its effect returns an outcome like HALTED_BEFORE_EFFECT or RECONCILIATION_REQUIRED instead of success. Every other project here reports what the model said it did. This one refuses to.
That is the same control the rest of this field keeps rediscovering from the other end — verify the side effect, not the summary — and it is why the record-and-replay layer is worth keeping on your map even though it cannot generalise. The trade is explicit and not subtle: a compiled flow does one task on one surface and breaks when the layout changes, which is why OpenAdapt declares its evidence sources per surface (DOM, accessibility, visual and OCR in browsers; UI Automation, Accessibility or AT-SPI on native desktops; pixels, OCR and anchors through RDP and Citrix). For the high-volume repetitive path — the 200-times-a-day form in the application with no API — paying a vision model to re-derive the same click sequence 200 times is a cost and a variance you chose for no reason.
When to pick which
| Situation | Pick | Why |
|---|---|---|
| Your agent clicks the wrong element, not the wrong step | UI-TARS | That is a grounding failure, and this is the only project here that ships weights you can pin, run locally and fine-tune |
| Your grounding is fine and the agent gives up or loops | Agent S3 | Planning, memory and retry are the policy layer; point it at UI-TARS for grounding as its README does |
| You need a hundred desktops for evaluation or training data | cua | Fleet VMs, four-language SDKs, bench harness and trajectory export; budget time for the per-component licences |
| One repetitive task, no API, high volume | OpenAdapt | Compile once, verify the effect every run, and stop paying a model to re-decide a known sequence |
| You are comparing them on an OSWorld number | Stop | Quote the benchmark version and the number of rollouts, or the figure means nothing; then measure on your own surfaces |
| Your shortlist includes Bytebot | Replace it | Archived read-only since 7 March 2026; cua covers the same need and is active |
FAQ
Is UI-TARS an alternative to Agent S3?
No. Agent S3 requires a grounding model and recommends UI-TARS-1.5-7B for the role, so the common configuration is both. If you see them benchmarked head-to-head, one of the two rows is measuring a stack that contains the other.
Should I use UI-TARS-2?
Not yet, from this repository. UI-TARS-2 was announced in September 2025, but the README's introduction, benchmark table and weights link all still refer to 1.5, and the page publishes no UI-TARS-2 weights or numbers. The pinnable artifact today is UI-TARS-1.5-7B.
Why is cua's licence described as "mostly MIT"?
Because the parts differ. The repository is MIT in the main, but Cua Spaces is FSL-1.1-MIT and becomes MIT two years after each release, cua-som is AGPL-3.0-or-later, and cua-perception is not MIT. Read the per-directory licences before a procurement review rather than the repository badge.
Does Behavior Best-of-N make Agent S3 better than a human?
It makes the number higher. 72.6% against a roughly 72% human baseline is a best-of-N result; the agent alone is 66%. The technique is legitimate, clearly labelled, and multiplies tokens, latency and — unless each attempt is isolated — side effects.
Which of these can I actually self-host end to end with no API key?
UI-TARS plus cua, with UI-TARS-1.5-7B served locally. Agent S3 needs a main generation model, which can be a vLLM endpoint but is usually hosted. OpenAdapt needs a model only while compiling, which is the strongest no-API-key story here for a task you can demonstrate.
What should I measure before choosing any of them?
Split your own failures into "wrong element" and "wrong step" over fifty trajectories. That one hour of labelling tells you which layer you are missing, and no published score can. See failure taxonomy and triage.
Further reading
On this wiki:
- GUI grounding — why "where to click" is a separate model with a separate failure mode.
- Computer use and GUI agents — the concept, from the top.
- Building computer-use agents — the playbook for shipping one.
- Human baselines in agent evals — what a best-of-N comparison against a person does and does not show.
- Screenshot and DOM artifacts — keeping the evidence a grounding failure needs.
- The benchmark landscape — why the version matters more than the score.
Project sources:
- UI-TARS — the model repository, paper and reported grounding benchmarks.
- UI-TARS Desktop / Agent TARS — the application shell around the model.
- Agent S — Agent S3, its model requirements and its benchmark table.
- cua — Spaces, Driver, Lume, CUA-S1 and Cua Bench, with the per-component licences.
- OpenAdapt — record, compile, replay, verify, and the effect-contract outcomes.
- OSWorld 2.0 — the 108-task long-horizon benchmark and its reported scores.