GUI Grounding

A44
Concepts · Agentic AI Explained

GUI grounding.

A computer-use agent that clicks 40 pixels left of the Save button and a computer-use agent that decides to delete the file instead of saving it have both failed your task, and your trace shows the same thing for both: a step that did not work. They need opposite fixes, and almost nobody separates them — which is why teams spend a quarter rewriting a planning prompt to repair a coordinate problem. Grounding is the component that turns "the Save control" into a point on a specific screen, and it is a different model, measured on different benchmarks, failing in a different way, from the one that chose to save.

STEP 1

The two decisions a GUI step contains.

Every action an agent takes on a screen resolves two questions in sequence. What should happen next — click Save, scroll to find the total, close the dialog that appeared — is a planning question answered from the task, the history and what the screen appears to contain. Where that action lands is a perception question answered from the pixels: which rectangle of this 1,512×982 screenshot is the Save control, and what are its coordinates.

In an API-driven agent loop the second question does not exist. A tool call names its target — save_document(id) — and the harness resolves it. On a screen there is no identifier, so something has to invent one. That something is the grounding model, and in most stacks it is literally a separate set of weights with a separate endpoint.

The giveaway that you have a grounding problem rather than a reasoning problem: the agent's own narration is correct. It says it is clicking Save, the trajectory records a click, and the file is not saved. A planning failure usually narrates the wrong intention; a grounding failure narrates the right one and misses.

STEP 2

Pixels, trees, and why the choice is usually made for you.

There are two ways to find a control, and they are not equally good — the reason both exist is that the better one is frequently unavailable.

  • Structured observation. The DOM in a browser, or the platform accessibility tree on a desktop — UI Automation on Windows, the Accessibility API on macOS, AT-SPI on Linux. Elements arrive with roles, names and bounding boxes already computed. Grounding becomes a lookup rather than a prediction, which makes it cheap, fast and nearly exact.
  • Visual grounding. A screenshot in, coordinates out, from a vision-language model trained for the job. Required for canvas-rendered applications, remote desktop and Citrix sessions, custom-drawn controls, native apps with no accessibility implementation, and anything behind a screen share. It is slower, costs a vision forward pass per step, and is the only option that works everywhere.

Prefer structure wherever it exists and fall back to pixels, which is what mature harnesses do per surface rather than per project. The cost of the fallback is not only accuracy: pixel grounding ties you to a resolution and a scale factor, so the same agent that works on your laptop misses on a 4K display at 150% scaling, and a retina screenshot downscaled before inference loses the 11-pixel icon entirely.

STEP 3

The benchmarks are separate, and that is the point.

Grounding has its own evaluations, and their existence is the cleanest evidence that it is its own component. ScreenSpot-style benchmarks hand a model one instruction and one screenshot and ask only for the coordinate — no planning, no multi-step execution, no environment. Task benchmarks such as OSWorld measure a whole agent completing real workflows, so a number from one tells you very little about the other.

The gap between the two is instructive. Published grounding accuracy on consumer interfaces runs into the 90s, while high-resolution professional-software grounding sits far lower and end-to-end task completion on long-horizon benchmarks lower still. Those are not inconsistent results; they are three different measurements, and the stack's ceiling is set by whichever is worst for your surfaces. If your application is a dense CAD or trading interface, the consumer-interface figure is irrelevant to you.

Label fifty of your own failed trajectories into "wrong element" and "wrong step" before you change anything. It takes about an hour, and it is the measurement that tells you whether to swap the grounding model, add structured observation for that surface, or rewrite the planner — three interventions that share no code. Without it, you are choosing by vibe between fixes that cost a quarter each.

STEP 4

Why grounding failures are disproportionately dangerous.

A reasoning failure usually produces a refusal, a loop or a visibly wrong plan. A grounding failure produces a confident action on the wrong object, which is the shape that gets past review.

  • Off-by-one-row. The agent meant the third invoice and clicked the fourth. Everything downstream is internally consistent and refers to the wrong entity, and nothing in the trace flags it.
  • The moved button. A dialog shifted 30 pixels between the screenshot and the click — a late-loading banner, a notification, an animation that had not settled. The agent clicked where the control was. See time-of-check to time-of-use; a screen is a mutable observation, and the window between seeing and acting is real.
  • Adjacent-destructive. Interfaces routinely place Archive beside Delete and Save beside Discard. A grounding error of one control width is the difference between a no-op and an irreversible action, which is why blast radius arguments for GUI agents have to assume the click lands next door.
  • Silent coordinate drift. A theme change, a browser-zoom setting or an OS update shifts everything by a few pixels. Accuracy degrades gradually with no error and no deploy to blame.

The structural response is to stop trusting the click as evidence that the action happened. Verify the effect through a channel the agent did not use to perform it — re-read the record, check the API, diff the state — and treat an unverifiable step as a failure rather than a success, which is the same discipline that defeats failure concealment for a completely different reason.

STEP 5

What this changes about buying and building.

Treat grounding as a procurable part with its own requirements, because it is one.

  • It is the piece you can pin. Open grounding models ship as weights under permissive licences, so this is the one layer of a computer-use stack whose behaviour you can freeze, run locally and fine-tune on your own interfaces — which matters most for exactly the dense professional software where general models are weakest.
  • It is also the piece a vendor hides. A hosted computer-use API gives you one score for the whole loop. When it regresses you cannot tell whether planning or grounding moved, and you cannot fix either.
  • Keep the screenshot. A grounding failure is unarguable from the screenshot plus the predicted coordinate, and uninvestigable without them. Retaining both per step is what makes the hour of labelling in STEP 3 possible at all.
  • Compile the repetitive path. If the same sequence runs hundreds of times a day on an unchanging surface, re-deriving the coordinates on every run buys variance and a bill. Record it once and replay a fixed sequence, with an effect check.

One concrete default: for any surface that exposes a DOM or an accessibility tree, ground against the tree and use the vision model only to disambiguate. You will give up the elegance of a single pixels-only pipeline and get back most of your failed steps, because the majority of production grounding errors are on controls that were fully described in a structure nobody read.