Mobile & Native App Agents

9 min read

U24
Playbook · Coding & Computer-Use Agents

Mobile & Native App Agents.

A coding agent's advantage is not that it writes better code than you — it is that it can be wrong twenty times an hour and keep the attempt that passed. Move that loop onto an iOS or Android codebase and a clean build alone eats minutes, a simulator has to boot, and the test that would have caught the regression renders a screen instead of returning a boolean. The answer is not a faster build. It is to restructure the target so the agent spends its attempts where compilation costs seconds, and to demote the full app build from the inner loop to a gate that runs once per candidate.

STEP 1

Budget the verify step first — everything else follows from it.

The localize-edit-verify loop is priced per attempt, and on native mobile the verify step is one to three orders of magnitude more expensive than on a backend service. Before you tune a prompt, measure the four numbers that set the ceiling on how good your agent can possibly be:

  • Incremental build of the module under edit — the number you want in seconds. On a well-modularised project it is; on a single-target app with a large bridging surface it is not, because touching one file invalidates half the graph.
  • Clean build of the whole app — minutes, and the number that decides whether the agent can afford a full build at all. Anything that forces a clean build (a dependency change, a generated-code step, a Gradle configuration edit) is a different class of task with a different budget.
  • Time to a running test. A JVM or Swift unit test starts almost immediately. An instrumented Android test or an XCUITest needs a booted emulator or simulator, an installed build, and an app launch — and the boot is often the largest single term.
  • Flake rate of the suite the agent will be graded on. This is the one teams skip, and it is the one that decides whether the loop converges at all. See Step 3.

Turn those four into one figure: verified attempts per hour. A backend agent gets dozens. An unmodified mobile app project frequently yields three or four, and at three attempts an hour the agent is not iterating — it is submitting drafts. Every recommendation below is an attempt to move that number, and you should re-measure it after each one rather than trusting that the change helped.

STEP 2

Split the target into a fast core and a slow shell.

The single highest-return change is not to the agent. It is to the codebase: make the part the agent edits compile and test without building the app. This is the same modularisation argument mobile teams have been having with themselves for a decade, with a new and unusually concrete payoff.

  • Pull logic out of view types. Networking, parsing, persistence, formatting, state reduction and business rules belong in modules with no UIKit/SwiftUI or Android framework dependency. Those modules build in seconds and test on the host machine with no device at all — that is where you want the agent to live.
  • Make the module boundary the agent's boundary. Scope the task to one module and tell the harness which build and test commands correspond to it. An agent that runs the whole suite because it does not know which target it changed has thrown away the modularisation you just paid for.
  • Keep a fast path even when the change is in the UI. A view-model change can be verified against the fast core; only the final candidate needs the shell. The distinction to hold is between iterating and confirming, and they do not need the same fidelity.
  • Hermetic build systems help, but do not wait for one. If you already run Bazel or an equivalent with remote caching, give the agent the cache — the win is large. If you do not, adopting one is a quarter of platform work, not a prerequisite for shipping an agent this month.

Where the codebase resists this — a single monolithic target, heavy code generation, a bridging header everything depends on — the honest read is that the agent will be weak there and the modularisation is the actual work. Say so in the plan rather than discovering it in the retro. Localisation across a large tree is its own problem; see repo navigation and code context.

STEP 3

Pin the device, or every failure is ambiguous.

An agent cannot tell a real regression from a flake. It can only tell that a test went red, and its trained response to red is to change code until it goes green — so a flaky suite does not slow the agent down, it actively teaches it to write wrong code. On a simulator or emulator the flake sources are enumerable, and every one of them is fixable configuration:

  • One pinned device definition — a named simulator or AVD with a fixed OS version, screen size and scale factor, created from a script that lives in the repo. Not "whatever was booted".
  • Fixed locale, region, time zone and clock. Date formatting and currency rendering are the classic overnight-failure pair, and a run that starts at 23:58 should behave like one that starts at noon.
  • Animations off, and a real idle signal. Disable UI animations at the system level and wait on an explicit idle condition rather than a sleep. Sleep-based waits are the largest single source of screen-test flake, and they get worse under the CPU contention of parallel agent runs.
  • Permissions, onboarding and auth pre-granted. Launch straight into a seeded, logged-in state with notification and location prompts already resolved. Every modal the agent has to dismiss is a failure mode you have volunteered for.
  • Deterministic network. Recorded fixtures or a local stub server, never the staging environment. A staging deploy mid-run turns a green suite red for reasons no diff explains.

Gate the agent on this: run your candidate suite twenty times against an unchanged commit and count the reds. If it is not zero, fix that before you grant the agent any autonomy on this codebase, because until it is zero you cannot distinguish "the agent broke it" from "the suite does that". This is the same discipline CI repair agents need, arriving earlier and mattering more.

STEP 4

Make snapshot tests the contract — and never let the agent update the references.

The agent cannot see your app. Whatever verification you give it is the entirety of its knowledge about whether the screen is right, and a passing unit test says nothing about a view that now renders behind the notch. Snapshot (screenshot) testing is the only mechanism that closes this gap at a cost the loop can afford: render the view to an image, compare against a committed reference, fail on a pixel difference above threshold.

  • Reference images are reviewed artefacts, committed to the repo, one set per device class and appearance mode you support. They are the specification of what the screen looks like, which makes them the thing a human reviewer should actually be reading in the diff.
  • The agent may propose a new reference; it may not accept one. This is the mobile form of the deleted-assertion failure that patch generation and test-driven loops warns about: re-recording the snapshot turns any red into a green, it is a single command, and it looks like progress in the transcript. Enforce it in CI — a diff that changes a reference image and is authored by the agent requires a human approval, mechanically, not by convention.
  • Test the views, not the flows, wherever you can. A snapshot of a view in a fixed state is fast and stable; a UI test that drives four screens to reach that state is slow and flaky. Push state in directly, render, compare.
  • Keep a small end-to-end suite anyway, and treat it as the gate, not the loop — it catches the wiring mistakes that per-view rendering cannot see. Ten scenarios that run once per candidate beat two hundred the agent runs never.
  • Attach the failing diff image to the agent's context. A vision-capable model given the before/after/difference triptych can often name the cause; given only "snapshot mismatch, 0.7%" it will guess. This is the one place a screenshot genuinely earns its tokens, and it is very different from driving the app through GUI control, which you should not be doing on your own codebase.
STEP 5

Draw the line at signing, entitlements and generated project files.

Native platforms have a release surface with no test that fails first — mistakes there surface as a rejected build, a broken update path, or a production crash on one OS version. Put these outside the agent's reach and say why:

  • Signing, provisioning profiles, keystores and entitlements. Credentials belong nowhere near an agent sandbox, and an entitlement change alters what the app is permitted to do at runtime — a security decision wearing a config file's clothes. Related: secrets management for agents.
  • Store metadata, privacy manifests and data-safety declarations. These are legal statements about your product. An agent may draft one; a human files it.
  • Generated project files. An agent that edits an Xcode project.pbxproj or regenerates a Gradle lockfile produces a diff no reviewer can read, which means it ships unreviewed. Either drive project structure from a declarative generator whose input is reviewable, or make these files agent-read-only.
  • Minimum-OS and dependency-version bumps. Small diffs with a large radius, and the compiler will not tell you which behaviour changed on the old OS you still support. Route them through dependency upgrade agents, which exist because this class needs its own protocol.
  • Anything that only fails on a physical device. Camera, Bluetooth, background execution, push delivery, battery behaviour. The simulator's silence here is not evidence.

The sandbox this all runs in deserves its own attention — a mobile build needs a large toolchain, a writable cache, and network access to a package registry, which is a broader grant than a typical agent job. See sandboxing and execution.

STEP 6

Pick the work by build cost per attempt.

Rank candidate tasks on one axis — how many full builds an attempt requires — and the backlog orders itself. The tasks that pay are the ones where the change is wide, the verification is cheap, and the review is mechanical:

  • Deprecation migrations across many call sites. Mechanical, compiler-verified, tedious for humans, and the diff reviews quickly because every hunk looks the same.
  • Localisation and string extraction. Hundreds of small edits, verified by a lint rule and a snapshot pass — and the pseudo-localised snapshot catches the truncation a translator would have reported three weeks later.
  • Test backfill on the fast core. No app build at all, so attempts are cheap and the artefact compounds. This is usually the best first project on a mobile codebase; see test generation agents.
  • Crash triage from symbolicated reports. The stack trace localises for you, which removes the expensive half of the loop. Pair with debugging and triage agents.
  • Accessibility labelling. Auditable by a rule, invisible to snapshot diffs, and chronically under-resourced — a rare case where the agent's tirelessness is the whole value.

What does not pay yet: new feature work in a heavily UI-coupled area, anything requiring judgement about motion or feel, and performance work whose signal only appears on a physical device under thermal load. Measure it rather than assuming — evaluating coding agents covers building the task set that tells you which column a given kind of work belongs in.

Before you write a single prompt, measure verified attempts per hour on your real repository and run your candidate suite twenty times on an unchanged commit. Those two numbers decide the project. If attempts per hour is under five, the work in front of you is modularisation and a pinned device — not agent configuration — and doing it in that order is the difference between an agent that lands merged pull requests and one that produces plausible drafts your team ends up rewriting. Then start on test backfill in the fast core, where the loop is already cheap and the artefact makes every later task cheaper.