AI Blog

UI-TARS vs Agent S3 vs cua vs OpenAdapt: you are picking a layer, not an agent

Agent S3 recommends UI-TARS as its grounding model, so two of these four are halves of one stack — and five numbers published under the name “OSWorld” span 63 points. Pick on which of four layers you are missing, and on the one axis that actually separates them: whether a model decides where to click every time the task runs.

By Agentic AI Wiki 13 min read

Agent S3's README recommends UI-TARS-1.5-7B as its grounding model, which means two of the four most-starred open-source computer-use projects are not competitors — they are the top and bottom halves of one stack. Read the other two the same way and the comparison stops being a leaderboard: you are choosing which of four layers you are missing, and the only axis that genuinely separates them is whether a generative model decides where to click every single time the task runs.

At a glance

Four projects, four layers, and one of them is the only one that ships the part that decides where the pointer goes. Figures read from each repository on 9 October 2026.

ProjectLayer it ownsLicenceShips the deciding part?
UI-TARS (ByteDance) The grounding model Apache-2.0 Yes — weights, on Hugging Face
Agent S3 (Simular) The policy around a model Apache-2.0 No — requires a main model and a grounding model
cua (trycua) The machine, and fleets of them Mostly MIT; see below Partly — CUA-S1 small models for narrow decisions
OpenAdapt A compiler from demonstration to program MIT Only at compile time — none at run time
Scores published under the name OSWorld Horizontal bars for five figures all reported as OSWorld results: 83.5 percent on OSWorld-Verified, 72.6 percent for Agent S3 with Behavior Best-of-N, 66 percent for Agent S3 alone, 44.33 percent for a leaderboard entry on OSWorld 2.0, and 20.6 percent for the best agent in the OSWorld 2.0 paper. Task success reported as “OSWorld” (%) 20 40 60 80 100 OSWorld-Verified Claude Opus 4.8, via a citing paper 83.5 OSWorld, 100 steps Agent S3 + Behavior Best-of-N 72.6 OSWorld, 100 steps Agent S3 alone 66.0 OSWorld 2.0, leaderboard Claude Opus 5, max effort — secondary source 44.3 OSWorld 2.0, paper best agent, binary completion at 500 steps 20.6
Five numbers, one benchmark name, a 63-point spread. None of them is wrong and none of them is comparable.

None of these four is a peer of another

The four layers of a computer-use agent, and who ships each A task enters a policy layer that plans and retries, which calls a grounding model that turns a screenshot into a coordinate, which acts on a machine holding the desktop. A side path shows a compiler that records one demonstration and then drives the machine directly with no model in the run-time loop. Every run Policy — what to do next, and after a failure plan, memory, retry, subtask decomposition Agent S3 (needs a main model) Grounding — where on this screen screenshot in, coordinate or element out UI-TARS (Apache-2.0 weights) The machine — pixels, files, clipboard VM or container desktop, driver, fleet, bench cua · Bytebot (archived 7 Mar 2026) The application under test a real GUI with no API for what you need Once, then never again Record one demonstration a human does the task; the compiler mines the effect contract from the observed delta OpenAdapt (MIT) Compiled program deterministic action sequence, per-surface evidence: DOM, accessibility, OCR, anchors Verify the effect, then report VERIFIED only if an independent check agrees; otherwise HALTED_BEFORE_EFFECT no generative model in the run-time loop
Three of the four stack on each other. The fourth replaces the top two after one run.

A computer-use agent has three jobs that are usually described as one. Something decides what to do next and what to do when it fails: that is the policy, and it is mostly prompt engineering, planning and retry logic. Something decides where on a 1,440-pixel-wide screenshot the "Save" control is: that is grounding, and it is a separate model with a separate failure mode. And something owns the pixels, the filesystem and the clipboard: that is the machine.

Sort the four projects by that split and the comparison resolves itself. UI-TARS is a vision-language model trained end-to-end for the second job, published as Apache-2.0 weights; its repository reports 94.2 on ScreenSpot-V2 and 61.6 on ScreenSpot-Pro for the 1.5 generation, which are grounding benchmarks, not task benchmarks. Agent S3 is the first job, explicitly: its setup requires a main generation model and separately requires a grounding model, and the one it recommends is UI-TARS-1.5-7B. cua is the third job, at fleet scale, with a bench harness and trajectory export attached. OpenAdapt is none of the three — it is a fourth thing we will get to.

This is why the star counts mislead. UI-TARS Desktop carries 39.2k stars and cua 29.1k, against 12.6k for Agent S3 and 1.8k for OpenAdapt — but UI-TARS Desktop is an application shell around the model, and the model repository itself sits at 11.6k. Popularity here tracks how many people wanted to try a desktop agent, not how much of the problem each project solves.

Why the published scores cannot pick for you

Two independent things are moving inside the numbers in that chart, and both of them are larger than the gap between any two of these projects.

The first is which benchmark the word "OSWorld" refers to. OSWorld-Verified is near saturation — a paper citing it puts Claude Opus 4.8 at 83.5% and reads that as desktop computer use being largely solved. OSWorld 2.0, posted in July 2026 with 108 long-horizon tasks built around exactly the phenomena the earlier set underrepresented, scores the same model family at 20.6% binary completion with a 54.8% partial score at 500 steps, and has GPT-5.5 plateauing near 13%. A 63-point difference under one name is not a disagreement about agents; it is a benchmark generation boundary, and a score quoted without the version is unusable. See the benchmark landscape for how routinely this happens.

The second is the selection budget. Agent S3's headline is 72.6% on OSWorld at 100 steps, which its README notes surpasses human performance of roughly 72%. The same README reports 66% for the agent alone; the extra 6.6 points come from Behavior Best-of-N, which runs the task several times and picks. The pattern repeats across its other two benchmarks: WindowsAgentArena 50.2% rising to 56.6% when selecting from three rollouts, AndroidWorld 68.1% rising to 71.6%. That is an honest and clearly-labelled technique, and it is also a cost multiplier rather than a capability: three rollouts is three times the tokens, three times the wall clock, and three times the side effects unless every attempt is sandboxed. Compare it against a human baseline and you are comparing N tries to one — which is the whole subject of human baselines in agent evals.

Strip both effects out and what remains is too small to choose on. Pick on the layer instead.

The layer everybody needs and nobody wants to maintain

What each project ships, and what it does at run time Matrix of UI-TARS, Agent S3, cua and OpenAdapt against five columns: ships weights, ships the policy, ships the machine, keeps a generative model in the run-time loop, and verifies the effect before reporting success. Only UI-TARS is strong on weights and only OpenAdapt on effect verification. What ships, and what runs Modelweights Thepolicy Themachine Model in therun-time loop Verifies theeffect UI-TARS Apache-2.0 No App shell only Every step No Agent S3 No Apache-2.0 No Two models No cua CUA-S1 only No Fleet + bench Whatever you add No OpenAdapt No Compiled No None Effect contract
Read down the last column. It is the only one where three of four are blank.

The machine is the unglamorous layer and the one that decides whether any of this reaches production, because an agent that drives a GUI needs a GUI, and a GUI that is also your laptop is not a deployment target. cua is the serious answer: VM-backed desktops on macOS, Linux and Omarchy, a driver that reaches native apps and browsers on macOS, Windows and Linux, Lume for local macOS and Linux VMs on Apple silicon, SDKs in four languages, and Cua Bench for building tasks and exporting trajectories for training. If you are evaluating or training rather than deploying one assistant, this is the project doing work nobody else here is doing.

Two caveats, both material. First, "MIT" is the repository's headline and not the whole licence: Cua Spaces — the desktop app, the part you would actually hand to users — is source-available under FSL-1.1-MIT and converts to MIT two years after each release, cua-som is AGPL-3.0-or-later, and the optional cua-perception extension is not MIT either. That is a reasonable structure, and it is also exactly the kind of detail that invalidates a licence-based procurement sign-off made from the badge. Second, neither cua nor any other project here publishes conventional GitHub releases, so pinning a version means pinning a commit or a component version string.

The cautionary case is the one most listicles still recommend. Bytebot — a self-hosted agent in a containerised Ubuntu 22.04 XFCE desktop with Firefox and VS Code, Apache-2.0, 11.1k stars, Helm and Compose deployment — was archived by its owner on 7 March 2026 and is read-only. It is a clean design and it is not a maintained dependency. If your shortlist came from a comparison post, check the archive banner before the feature table; this is the second category in a week where the most-linked project turned out to be frozen.

One of them takes the model out of the loop

OpenAdapt is the smallest project here by an order of magnitude and the only one that answers a different question. You demonstrate the task once; openadapt flow record captures it, openadapt flow compile turns the recording into a bundle, and the replay runs the compiled program without a generative-model API in the control loop. The compiler's job, in the project's own phrasing, is to mine the effect contract from the observed delta — it works out what the demonstration actually changed, and that becomes the thing the run is judged against.

Which produces the property that makes it worth your attention regardless of whether you adopt it: it reports VERIFIED only if an independent check agrees. The tutorial confirms a saved record through a read-only API that the writing screen never touches, and a run that cannot establish its effect returns an outcome like HALTED_BEFORE_EFFECT or RECONCILIATION_REQUIRED instead of success. Every other project here reports what the model said it did. This one refuses to.

That is the same control the rest of this field keeps rediscovering from the other end — verify the side effect, not the summary — and it is why the record-and-replay layer is worth keeping on your map even though it cannot generalise. The trade is explicit and not subtle: a compiled flow does one task on one surface and breaks when the layout changes, which is why OpenAdapt declares its evidence sources per surface (DOM, accessibility, visual and OCR in browsers; UI Automation, Accessibility or AT-SPI on native desktops; pixels, OCR and anchors through RDP and Citrix). For the high-volume repetitive path — the 200-times-a-day form in the application with no API — paying a vision model to re-derive the same click sequence 200 times is a cost and a variance you chose for no reason.

When to pick which

SituationPickWhy
Your agent clicks the wrong element, not the wrong step UI-TARS That is a grounding failure, and this is the only project here that ships weights you can pin, run locally and fine-tune
Your grounding is fine and the agent gives up or loops Agent S3 Planning, memory and retry are the policy layer; point it at UI-TARS for grounding as its README does
You need a hundred desktops for evaluation or training data cua Fleet VMs, four-language SDKs, bench harness and trajectory export; budget time for the per-component licences
One repetitive task, no API, high volume OpenAdapt Compile once, verify the effect every run, and stop paying a model to re-decide a known sequence
You are comparing them on an OSWorld number Stop Quote the benchmark version and the number of rollouts, or the figure means nothing; then measure on your own surfaces
Your shortlist includes Bytebot Replace it Archived read-only since 7 March 2026; cua covers the same need and is active

FAQ

Is UI-TARS an alternative to Agent S3?

No. Agent S3 requires a grounding model and recommends UI-TARS-1.5-7B for the role, so the common configuration is both. If you see them benchmarked head-to-head, one of the two rows is measuring a stack that contains the other.

Should I use UI-TARS-2?

Not yet, from this repository. UI-TARS-2 was announced in September 2025, but the README's introduction, benchmark table and weights link all still refer to 1.5, and the page publishes no UI-TARS-2 weights or numbers. The pinnable artifact today is UI-TARS-1.5-7B.

Why is cua's licence described as "mostly MIT"?

Because the parts differ. The repository is MIT in the main, but Cua Spaces is FSL-1.1-MIT and becomes MIT two years after each release, cua-som is AGPL-3.0-or-later, and cua-perception is not MIT. Read the per-directory licences before a procurement review rather than the repository badge.

Does Behavior Best-of-N make Agent S3 better than a human?

It makes the number higher. 72.6% against a roughly 72% human baseline is a best-of-N result; the agent alone is 66%. The technique is legitimate, clearly labelled, and multiplies tokens, latency and — unless each attempt is isolated — side effects.

Which of these can I actually self-host end to end with no API key?

UI-TARS plus cua, with UI-TARS-1.5-7B served locally. Agent S3 needs a main generation model, which can be a vLLM endpoint but is usually hosted. OpenAdapt needs a model only while compiling, which is the strongest no-API-key story here for a task you can demonstrate.

What should I measure before choosing any of them?

Split your own failures into "wrong element" and "wrong step" over fifty trajectories. That one hour of labelling tells you which layer you are missing, and no published score can. See failure taxonomy and triage.

Further reading

On this wiki:

Project sources:

  • UI-TARS — the model repository, paper and reported grounding benchmarks.
  • UI-TARS Desktop / Agent TARS — the application shell around the model.
  • Agent S — Agent S3, its model requirements and its benchmark table.
  • cua — Spaces, Driver, Lume, CUA-S1 and Cua Bench, with the per-component licences.
  • OpenAdapt — record, compile, replay, verify, and the effect-contract outcomes.
  • OSWorld 2.0 — the 108-task long-horizon benchmark and its reported scores.