Playbooks / Coding & Computer-Use Agents

Coding & Computer-Use Agents

Agents that read code, write code, run tools, and drive a computer — patterns, harnesses, and pitfalls.

  1. Coding Agent Architecture
    The localize-edit-verify loop that makes a coding agent more than a code generator: the agent-computer interface, why agentic beats pipeline coding, and where the loop fails.
  2. Repo Navigation & Code Context
    Code search vs. embeddings, symbol-level indexing, context budgeting over a large tree, and why confident wrong localization is the expensive failure of code retrieval.
  3. Patch Generation & Test-Driven Loops
    Structured diffs and hunk-apply failures, test-driven self-correction, regression guarding, and the three honest liars in the loop: flakes, overfit, and the deleted assertion.
  4. Computer-Use & GUI Agents
    Pixel vs. DOM grounding, the action space, the screenshot loop, and the multiplicative latency and reliability tax that makes GUI control a last resort.
  5. Browser agents
    Driving a real browser as a tool — DOM versus pixel observation, login + auth state, the well-trodden failure modes, and when to step up to a full GUI agent.
  6. IDE agents
    Coding agents that live in the editor — the loop is the same as a CLI coding agent, but the interaction surface, undo expectations, and trust threshold are all different.
  7. Sandboxing & Safe Execution
    Containerized execution, network and filesystem isolation, capability scoping, and designing for blast radius when an agent runs untrusted, attacker-influenced code.
  8. Evaluating Coding Agents
    The SWE-bench family, pass@k vs. resolve rate, harness sensitivity, documented contamination, and why a private post-cutoff eval set is the only number to trust.
  9. Code Review Agents
    A review bot lives or dies on precision, not recall: diff-anchored context, an adversarial gate that drops any finding without a concrete failure scenario, a hard comment budget ranked worst-first, and acted-upon rate as the one production metric.
  10. Large-Scale Migration Agents
    Generation went to zero and human review did not, so the deliverable is evidence rather than patches: build the oracle first, batch by verifiability instead of by directory, hand the mechanical head to a codemod and only the tail to the model, and let a hundred-file pilot’s unedited-merge rate decide whether the project is viable.
  11. Debugging & Triage Agents
    An agent that reads a stack trace and emits a diff has pattern-matched, not debugged — make the failing test the deliverable and the eval criterion becomes objective, the spend moves from generation to observation, and "cannot reproduce" becomes a result you can trust.
  12. Background Coding Agents
    An agent that opens twelve pull requests a day adds nothing if your team merges four — detaching from the editor moves the bottleneck to review, so every decision is either about making a run self-verifying or about keeping the queue short enough that the work lands.
  13. Test-Generation Agents
    An agent that writes tests from your code infers the spec from the implementation, so wherever the code is wrong the test now certifies the bug and blocks the fix — coverage cannot see this, mutation score can, and the only jobs worth dispatching are the ones where you can name the oracle in a sentence.
  14. Dependency Upgrade Agents
    Bumping a version has been automated since 2017 and is worth nothing — the backlog exists because nobody will merge an upgrade they cannot vouch for, so what you are building is an evidence policy: a green suite proves least about the changed defaults that break silently, and correct triage beats upgrades merged.
  15. Documentation Agents
    Documentation is the one artefact in the repo with no oracle, and a wrong paragraph gets believed for years — so split the corpus by what a machine can falsify, give the agent unsupervised authority over derived reference and executable prose only, make deletion a first-class output, and report freshness rather than pages written.
  16. Design-to-Code Agents
    A screenshot cannot say "this is the Button component", so a pixel-faithful generator rebuilds one, hard-codes the hex instead of the token, and passes every visual review — judge the agent on component reuse rate and token adherence instead, hand it the typed component API and the design-to-import mapping before the picture, and if a ten-screen pilot comes in under seventy per cent reuse you have a design-system problem the agent will only scale.
  17. Vulnerability Remediation Agents
    Generating the patch is the cheap half — about a quarter of model-written patches for post-cutoff CVEs fix the bug without changing behaviour, and more than forty per cent of the ones your validation calls correct fail once you test mutated exploits rather than the one you were handed, so the deliverable is a reproduction that fails before and passes after, on a queue triaged by reachability instead of scanner severity.
  18. Spec-Driven Development with Coding Agents
    The plan, the tests, the code and the summary all descend from one reading of the task, so when that reading is wrong every artefact agrees and the review passes — a spec earns its place only insofar as it comes from outside that loop, which means human-authored observable criteria with an explicit non-goals list, a check id or a named human beside each one, and spec and behaviour changing in the same pull request.
  19. Performance-Optimization Agents
    Every other coding agent gets a verifier that cannot lie; this one gets a stopwatch — so forty candidate patches scored on one timing run each will keep noise with near-certainty. Build the harness first (confidence intervals, a published minimum detectable effect, a fixed repetition budget), gate on the full test suite rather than folding correctness into the score, and hand the agent a profile instead of a repository.
  20. Infrastructure-as-Code Agents
    Every other coding agent has to build its own oracle; this one is handed terraform plan for free — so the deliverable is a plan whose every line falls into a class you already agreed to auto-apply, the review job becomes classifying the plan rather than reading the diff, and the engineering is all in the four things a plan is silent about: unknown values, server-side behaviour, replacements announced in the same tone as everything else, and whatever is not in state.
  21. CI Repair Agents
    Handed a reproduction on a plate, this agent’s real job is the classification nobody runs: caused by this diff, pre-existing on the base, infrastructure, or non-deterministic — prove causation by running the failing check against the merge base before pushing, spend at most one re-run, never touch what a test asserts, and page yourself on wrong-push rate rather than builds turned green.
  22. Database-Migration Agents
    A model writes correct DDL first time, which is why this is the coding-agent task most likely to take production down: the statement is fine and the sequence is wrong, and CI runs it against an empty table with nobody else connected. Make the deliverable an ordered expand–backfill–contract sequence the currently-deployed code still fits, default the agent to additive-only changes, emit the lock_timeout preamble every time, and grade it on lock duration measured by a shadow apply against a production-shaped clone under load — the agent authors, a controlled runner applies.
  23. Notebook & Data-Science Agents
    The file on disk is not the program that produced the answer — the kernel is, and nothing writes it down, so “it ran” is unfalsifiable until a cold restart-and-run-all executed by your harness says otherwise. Make that gate the definition of done and hand it to the agent as a tool, feed it schema cards instead of printed dataframes, remember the dangerous boundary is the warehouse credential rather than the sandbox, and grade the number and the method rather than whether the code ran.
  24. Mobile & Native App Agents
    A coding agent’s advantage is being wrong twenty times an hour, and a clean iOS or Android build spends that budget before lunch — so the fix is not a faster build but a codebase split into a fast core the agent iterates in and a slow shell the full build gates once per candidate. Pin the simulator or every red is ambiguous, make committed snapshot references the contract the agent may propose but never accept, and rank the backlog by full builds per attempt.
  25. Accessibility Remediation Agents
    Give an agent a scanner score and it builds you an accessibility overlay inside your own repository, because the cheapest way to silence a rule is an ARIA attribute that lies — and automated testing reaches roughly half the failures by count and almost none of the ones that block a journey. Make the unit of work a keyboard-only user journey, fix at the design system rather than the call site, and require every pull request to carry a before-and-after focus trace as its proof.
  26. Dead-Code & Feature-Flag Removal Agents
    Deletion is the one coding-agent task where a green suite proves nothing — it passes for exactly the reason the code looked dead — and coverage rises when you delete untested code, so the obvious objective rewards removing what you understand least. Static analysis proposes and never decides; the evidence has to be runtime reachability over a window set by the business calendar, not a round thirty days. A flag evaluated a million times returning false is not a flag never evaluated, a kill switch reads exactly like a stale flag to every heuristic, and the production value is usually the opposite of the code default.
  27. Merge Queues for Agent-Authored Changes
    Review is not the bottleneck — 31% more PRs now merge with no review at all, and the incidents-to-PR ratio tripled, so the system absorbed the volume by routing around its own gate. The serialized path to main is the gate that is left, and batching (the standard fix) inverts under agent load: at a 20% per-PR queue-failure rate a batch of ten passes 10.7% of the time and costs you a full CI run per merge. Size batches to measured p, insist on bisection, and rebase-and-verify against the queue head before anything enters.
  28. Release & Publishing Agents
    Every other coding-agent task is revertible; a published release is not — npm’s unpublish window is 72 hours, the version number can never be reused, and the lockfiles that already resolved it are permanent. So the design centre is the signing boundary, not the automation: let the agent prepare the bump, the notes and the dry run, and arrange that no long-lived publish credential exists for it to hold. Provenance attests where an artefact was built, not that a human agreed to ship it — so if your agent can push to the ref your trusted publisher builds from, the attestation is valid and says the wrong thing. Measure time-to-supersede, and make the forward-only fix path a tool the agent can call.
  29. Secret-Scanning & Rotation Agents
    Detection is the solved half and you are paying for it twice — over 64% of credentials confirmed valid in 2022 were still valid at a 2026 retest, so the deliverable is a credential proven dead, with a receipt. Proving it requires using it, which forces the architecture: the agent sees fingerprints while a separate verifier service holds the secret, and the agent may create credentials but never revoke them.