Playbooks / Domain Playbooks

Domain Playbooks

Domain-specific playbooks — customer support, research, sales, data analysis, DevOps — what to build, what to skip.

  1. Customer-Support Agents
    Optimize deflection subject to a near-zero confident-wrong-answer rate: grounded answers autonomous, transactions gated, tone graded in the eval, clean handoff over a confident guess.
  2. Data & Analytics Agents
    The failure mode is a confidently wrong number: schema/semantic-layer grounding, read-only execution, verification as a separate stage, and abstention scored above confident error.
  3. DevOps & SRE Agents
    Read-only first because the blast radius is production: diagnosis before remediation, runbooks as tested tools, limits in the tool signature, change control unchanged by the operator being a model.
  4. Research & Synthesis Agents
    A fabricated source voids the whole deliverable: retrieve-then-write-then-verify, citation faithfulness as a hard constraint, disagreement preserved not averaged, depth vs breadth as a bounded budget.
  5. Sales & GTM Agents
    The failure mode is automated spam at scale: value in research/personalization not volume, consent as an upstream fail-closed gate, human sign-off scaling with reach, spam-risk weighted heavily in the eval.
  6. Finance agents
    Where agents earn their keep in finance — reconciliation, research synthesis, KYC review — and the hard rails (audit, determinism, regulator-readable trails) they must carry.
  7. Healthcare agents
    Charting, prior auth, intake triage — the few healthcare jobs where agents shave real labor, and the privacy + clinical-safety guardrails you cannot ship without.
  8. Legal agents
    Discovery, contract review, citation checking — where legal agents already work, where they hallucinate, and what supervision they need by jurisdiction.
  9. Adapting a Playbook to Your Domain
    The meta-method behind every playbook: derive a new vertical by answering five questions in order — job, autonomy by reversibility, tools as grounding-and-limit, eval mirroring the cost asymmetry, structural guardrails.
  10. Hiring & Recruiting Agents
    The one domain where regulators specified the architecture first: NYC Local Law 144 and EU AI Act Annex III demand a countable, attributable per-candidate decision — which is why the free-text "strong fit, 8/10" design fails an audit, and why the model belongs on the widening side of the funnel.
  11. Email & Calendar Agents
    The only system of record an unauthenticated stranger can write to: provenance tiers that survive a forward, a reader/actor split with a typed interface so message content can never reach the send path, calendar as the zero-click vector nobody hardens, and confirmations that ask only when something is unusual.
  12. Tutoring & Learning Agents
    The only domain where doing the task well is the failure: measure unaided post-test rather than session satisfaction, enforce the hint ladder in code, diagnose the misconception instead of explaining the topic, and accept that any agent able to do the homework has already broken homework as assessment.
  13. Security-Operations Agents
    The only agent whose input is authored by an adversary who knows a model reads it: deterministic enrichment with the verdict kept out of the model, autonomy by reversibility with containment gated, and a golden set built from the closures you were never told were wrong.
  14. Translation & Localization Agents
    Fluency stopped predicting fidelity, so reading the target text catches nothing: review the source-target pair in a diff, gate on machine-checkable invariants — placeholder parity, termbase, structure — and treat terminology as retrieval rather than a glossary in the prompt.
  15. Shopping & Checkout Agents
    Instant Checkout was pulled back with fewer than fifteen merchants live because product data, not payments, was the hard part — so ground the item, re-verify at purchase time, and gate the one irreversible step outside the model.
  16. Content Moderation Agents
    The published guidelines are a summary of a decade of unwritten precedent, so build moderation as retrieval over decided cases — and let the base rate, not the accuracy score, tell you how many reviewers you still need.
  17. Insurance Claims Agents
    Cycle time is spent waiting for a missing document, not deciding — so build the completeness engine, stop at the coverage determination, and decline fraud scoring deliberately: it is the use case with the worst risk-adjusted return in the domain.
  18. IT Helpdesk Agents
    The tickets that cost you all end in a mutation of an identity system, so the product is the authorisation layer: proof identity out of band on a possession factor, derive the action set from entitlements rather than the request, and ship password and MFA recovery last — or never.
  19. KYC & AML Onboarding Agents
    With 90–95% of screening alerts false positive, the constraint was never detection but the backlog of alerts nobody can write up defensibly — so the agent assembles evidence and drafts the rationale, a named human disposes, and the sanctions matcher stays deterministic and versioned.
  20. Procurement & Sourcing Agents
    A losing bidder is entitled to the reasons for the award, so the one thing to automate is not the score: build the requirement-coverage matrix with page-level locators, isolate each bid in its own index, and keep ranking with a named human who can defend it.
  21. Public-Benefits Casework Agents
    MiDAS auto-adjudicated fraud for 34,000 people at a roughly 93% error rate and the Dutch benefits scandal took down a cabinet — so the agent assembles the packet and a named human determines, retrieval is pinned to the rules in force on the claim date, and the error you must measure is the denial nobody appealed.
  22. Travel & Booking Agents
    Search is free and repeatable; the booking is a payment, a contract and a seat someone else loses — so build them as two systems, and make the committed path accept a quote ID rather than parameters, re-price at commit, carry an idempotency key and reconcile instead of retrying.
  23. Supply-Chain & Logistics Agents
    The failure mode is not hallucination, it is acting confidently on a fact that was true four hours ago — so put a mandatory as_of on every tool response with a code-enforced maximum age, leave routing to the solver and scope the agent to a closed exception taxonomy, and measure the exception it never raised rather than the precision of the ones it did.
  24. Field Service & Dispatch Agents
    Scheduling is the one part already solved — a constraint solver beats a model at assignment, so the agent belongs at the two edges it cannot read: intake that decides which parts go on the truck, and the reschedule call when the day breaks. Every board write is a promise, so reserve-then-confirm it, assume human dispatchers are writing too, and measure first-time-fix rather than automation rate.
  25. Accounts Payable & Invoice Agents
    Best-in-class touchless has hovered near 49% for years and the failing half is not failing on reading the document — it fails because there is no PO, no receipt, or nobody ordered it, so build for the exception queue rather than the clean lane. Gate on irreversibility instead of amount: a bank-detail change is the loss you never claw back.
  26. Collections & Dunning Agents
    Regulation F allows seven calls per debt per seven days; New York City's SHIELD rule allows three communications of any kind per account — so the contact governor, not the model, is the product, and every channel must decrement one authoritative counter before it composes a word. Right-party contact comes before content, and disclosures, balances and offer ladders belong to code.
  27. Marketing & Ad-Operations Agents
    The agent gets a live spend lever and a feedback number that is noise at the timescale it wants to act on, while the ad platform already runs its own optimiser on the same account — so this is a controller-design problem: a minimum change interval, an observation window measured in conversions rather than hours, an objective function taken from your warehouse rather than from the party you are paying, and a read-only reconciliation agent that recovers more money from broken tracking than any bidding change will.
  28. Fraud & Disputes Agents
    False declines cost the industry roughly 13× the fraud they prevent, and the transactions you decline never generate a label at all — so the only honest measurement is a budgeted approve-anyway holdout, the dashboard has to report cohorts old enough to be true against a 120-day dispute window, and the agent belongs in the case file rather than in a hundred-millisecond authorisation path.
  29. Medical Coding & Claims Agents
    The 835 remittance hands you a free adversarial label on every claim, and it is biased in exactly one direction — undercoding is never denied, so a single-sided objective has a degenerate optimum the optimiser will find; make the scorecard two-sided against a blind-coded baseline, route on the CARC/RARC pair because CO-16 is a pointer rather than a reason, and treat autonomy as a jurisdiction-keyed routing decision now that Indiana requires human review before submission.
  30. RFP & Security-Questionnaire Agents
    Every answer is a contractual representation, so a stale claim is a misrepresentation, not a hallucination: an answer library with an owner and a valid_until per claim, matching that scores the difference not the similarity, three tiers with a hard block on future-tense commitments.
  31. Expense & Travel Audit Agents
    The recoverable money is unreclaimed VAT, duplicates and priced-in policy leakage — not fraud — so scale the deterministic checks to 100% and leave the judgement calls sampled, keep policy in an effective-dated rule engine the model never reads, and score the one job the model owns (reconciling receipt, card feed and claim) against a card-settlement ground truth that arrives free every night.
  32. Clinical-Trial Matching Agents
    The expensive error is invisible — an eligible patient who was never surfaced — so build for recall, not for the screen-failure rate you can see: decompose free-text criteria into atomic predicates that each carry a time window, emit a per-criterion table with evidence spans and an explicit unknown state instead of a verdict, and gate patient contact behind a named human.
  33. Pharmacovigilance & Adverse-Event Agents
    Putting an agent on a channel is a legal act, not a monitoring decision: the reporting clock starts at first knowledge by anyone acting for the company, so the written channel scope comes before the model. Build around the four elements that make a case valid, emit missing-element follow-up questions instead of a reportable yes/no, constrain coding to the dictionary version in force, and let the agent escalate while only a qualified person may dismiss.
  34. Drive-Thru & Restaurant Ordering Agents
    Voice AI alone gets about 83% of drive-thru orders right against 87% for the standard lane, and about 95% when staff step in — which they do on roughly one order in five — so the product is the handoff, not the recogniser. Track containment, intervention cost, order time and drive-offs rather than accuracy; ground every item in the live menu version and write through the POS; and remember the hard part is the modifier grammar and the lane’s audio, not the transcription.
  35. Ambient Clinical Documentation Agents
    The headline result moved burnout from 51.9% to 38.8% and did not measure documentation quality at all — so time-to-signature, the metric every dashboard ships, is maximised by a clinician who signs without reading. Design backwards from the signature as an authorship transfer, enforce provenance tiers so a physical-exam finding appears only if it was said aloud, surface omissions because review cannot catch them, keep code assignment out of the scribe, and fund the weekly adjudicated sample that is the only number measuring the note.
  36. Prior Authorization Agents
    Statute has already split this workflow: across eleven states in two sessions an AI may not be the sole basis for a medical-necessity denial, while the same laws expressly permit AI for administrative work and organising clinical information — so build an agent that can approve and escalate but cannot deny, with the denial branch absent rather than flagged off. Since January 2026 payers owe 72 hours expedited and 7 days standard with a specific reason attached, so engineer against the reason code and first-pass completeness rather than an approval-probability model — and read WISeR, where Texas requests ran 62% approved by machine and 84% after human review, as a lesson about the objective rather than the technology.
  37. Insurance Underwriting Agents
    Bind rate arrives in seconds and loss experience arrives after the development tail, so any loop closed on the fast signal is an adverse-selection machine — and the regulated artifact is the enumerable factor that sets price, which an LLM judgement can never be. Emit declared variables with a source span and an explicit unknown, leave the rating engine deterministic, run the BIFSG-style outcome test on your own book before an examiner does, and make the decline path a referral rather than a branch.
  38. Smart-Home & IoT Agents
    Every demo turns on a light and every incident will be about something the agent read: device event history is a presence log for the whole household, attacker-writable at the device-name layer and impossible to un-read once it is in a context window. Classify devices by reversibility rather than category, write policy over runtime-discovered traits with default-deny, express commands as absolute targets verified by observation, cap the history window server-side, and design for the people in the building who never saw the consent screen.
  39. Consumer Credit Agents
    Teams brace for the explainability problem and get caught by the conversation: under Regulation B what the agent asks is restricted, what it says to a hesitant applicant can be unlawful discouragement, and the moment it stops collecting documents starts a thirty-day notice clock nobody wired a timer to. Build intake as a closed question registry with a coded completeness rule, keep scoring in a declared model the agent cannot reach, and measure document-extraction accuracy on the denial population — that tail is what produces wrong principal reasons.
  40. Payroll & Employment-Tax Agents
    Payroll already contains a deterministic authority, so an agent that produces a figure destined for a filing has inserted a probabilistic step into a process whose whole penalty structure assumes a determinate one. Keep it on the input and reading sides of the engine, and note that the primary control is a timer rather than a confidence threshold: US deposit penalties step 2/5/10/15% by lateness alone, so an agent that asks a question and waits has produced a compliance failure. Spend the capability on notice triage, variance and reconciliation, and never let a model infer one of 7,400+ local jurisdictions.
  41. Scientific-Discovery Agents
    Two benchmarks posted to arXiv in early October 2026 settle the design question: on EurekaBench an agent hit 47.4% predictive accuracy against a human scientist's 48.8% and 29.4% on the insight the task was built around against 69.7% — so build for the metric it already wins and you ship expert-accuracy correlations nobody can publish. Split execution, analysis and mechanism-proposal into separate agents, pre-run the expensive simulations and grade with fixed rules rather than a judge, declare guidance as a logged L0–L3 parameter, and report expert-accepted insights per expert review-hour.