Playbooks / Domain Playbooks
Domain Playbooks
Domain-specific playbooks — customer support, research, sales, data analysis, DevOps — what to build, what to skip.
- Customer-Support AgentsOptimize deflection subject to a near-zero confident-wrong-answer rate: grounded answers autonomous, transactions gated, tone graded in the eval, clean handoff over a confident guess.
- Data & Analytics AgentsThe failure mode is a confidently wrong number: schema/semantic-layer grounding, read-only execution, verification as a separate stage, and abstention scored above confident error.
- DevOps & SRE AgentsRead-only first because the blast radius is production: diagnosis before remediation, runbooks as tested tools, limits in the tool signature, change control unchanged by the operator being a model.
- Research & Synthesis AgentsA fabricated source voids the whole deliverable: retrieve-then-write-then-verify, citation faithfulness as a hard constraint, disagreement preserved not averaged, depth vs breadth as a bounded budget.
- Sales & GTM AgentsThe failure mode is automated spam at scale: value in research/personalization not volume, consent as an upstream fail-closed gate, human sign-off scaling with reach, spam-risk weighted heavily in the eval.
- Finance agentsWhere agents earn their keep in finance — reconciliation, research synthesis, KYC review — and the hard rails (audit, determinism, regulator-readable trails) they must carry.
- Healthcare agentsCharting, prior auth, intake triage — the few healthcare jobs where agents shave real labor, and the privacy + clinical-safety guardrails you cannot ship without.
- Legal agentsDiscovery, contract review, citation checking — where legal agents already work, where they hallucinate, and what supervision they need by jurisdiction.
- Adapting a Playbook to Your DomainThe meta-method behind every playbook: derive a new vertical by answering five questions in order — job, autonomy by reversibility, tools as grounding-and-limit, eval mirroring the cost asymmetry, structural guardrails.
- Hiring & Recruiting AgentsThe one domain where regulators specified the architecture first: NYC Local Law 144 and EU AI Act Annex III demand a countable, attributable per-candidate decision — which is why the free-text "strong fit, 8/10" design fails an audit, and why the model belongs on the widening side of the funnel.
- Email & Calendar AgentsThe only system of record an unauthenticated stranger can write to: provenance tiers that survive a forward, a reader/actor split with a typed interface so message content can never reach the send path, calendar as the zero-click vector nobody hardens, and confirmations that ask only when something is unusual.
- Tutoring & Learning AgentsThe only domain where doing the task well is the failure: measure unaided post-test rather than session satisfaction, enforce the hint ladder in code, diagnose the misconception instead of explaining the topic, and accept that any agent able to do the homework has already broken homework as assessment.
- Security-Operations AgentsThe only agent whose input is authored by an adversary who knows a model reads it: deterministic enrichment with the verdict kept out of the model, autonomy by reversibility with containment gated, and a golden set built from the closures you were never told were wrong.
- Translation & Localization AgentsFluency stopped predicting fidelity, so reading the target text catches nothing: review the source-target pair in a diff, gate on machine-checkable invariants — placeholder parity, termbase, structure — and treat terminology as retrieval rather than a glossary in the prompt.
- Shopping & Checkout AgentsInstant Checkout was pulled back with fewer than fifteen merchants live because product data, not payments, was the hard part — so ground the item, re-verify at purchase time, and gate the one irreversible step outside the model.
- Content Moderation AgentsThe published guidelines are a summary of a decade of unwritten precedent, so build moderation as retrieval over decided cases — and let the base rate, not the accuracy score, tell you how many reviewers you still need.
- Insurance Claims AgentsCycle time is spent waiting for a missing document, not deciding — so build the completeness engine, stop at the coverage determination, and decline fraud scoring deliberately: it is the use case with the worst risk-adjusted return in the domain.
- IT Helpdesk AgentsThe tickets that cost you all end in a mutation of an identity system, so the product is the authorisation layer: proof identity out of band on a possession factor, derive the action set from entitlements rather than the request, and ship password and MFA recovery last — or never.
- KYC & AML Onboarding AgentsWith 90–95% of screening alerts false positive, the constraint was never detection but the backlog of alerts nobody can write up defensibly — so the agent assembles evidence and drafts the rationale, a named human disposes, and the sanctions matcher stays deterministic and versioned.
- Procurement & Sourcing AgentsA losing bidder is entitled to the reasons for the award, so the one thing to automate is not the score: build the requirement-coverage matrix with page-level locators, isolate each bid in its own index, and keep ranking with a named human who can defend it.
- Public-Benefits Casework AgentsMiDAS auto-adjudicated fraud for 34,000 people at a roughly 93% error rate and the Dutch benefits scandal took down a cabinet — so the agent assembles the packet and a named human determines, retrieval is pinned to the rules in force on the claim date, and the error you must measure is the denial nobody appealed.
- Travel & Booking AgentsSearch is free and repeatable; the booking is a payment, a contract and a seat someone else loses — so build them as two systems, and make the committed path accept a quote ID rather than parameters, re-price at commit, carry an idempotency key and reconcile instead of retrying.
- Supply-Chain & Logistics AgentsThe failure mode is not hallucination, it is acting confidently on a fact that was true four hours ago — so put a mandatory as_of on every tool response with a code-enforced maximum age, leave routing to the solver and scope the agent to a closed exception taxonomy, and measure the exception it never raised rather than the precision of the ones it did.
- Field Service & Dispatch AgentsScheduling is the one part already solved — a constraint solver beats a model at assignment, so the agent belongs at the two edges it cannot read: intake that decides which parts go on the truck, and the reschedule call when the day breaks. Every board write is a promise, so reserve-then-confirm it, assume human dispatchers are writing too, and measure first-time-fix rather than automation rate.
- Accounts Payable & Invoice AgentsBest-in-class touchless has hovered near 49% for years and the failing half is not failing on reading the document — it fails because there is no PO, no receipt, or nobody ordered it, so build for the exception queue rather than the clean lane. Gate on irreversibility instead of amount: a bank-detail change is the loss you never claw back.
- Collections & Dunning AgentsRegulation F allows seven calls per debt per seven days; New York City's SHIELD rule allows three communications of any kind per account — so the contact governor, not the model, is the product, and every channel must decrement one authoritative counter before it composes a word. Right-party contact comes before content, and disclosures, balances and offer ladders belong to code.
- Marketing & Ad-Operations AgentsThe agent gets a live spend lever and a feedback number that is noise at the timescale it wants to act on, while the ad platform already runs its own optimiser on the same account — so this is a controller-design problem: a minimum change interval, an observation window measured in conversions rather than hours, an objective function taken from your warehouse rather than from the party you are paying, and a read-only reconciliation agent that recovers more money from broken tracking than any bidding change will.
- Fraud & Disputes AgentsFalse declines cost the industry roughly 13× the fraud they prevent, and the transactions you decline never generate a label at all — so the only honest measurement is a budgeted approve-anyway holdout, the dashboard has to report cohorts old enough to be true against a 120-day dispute window, and the agent belongs in the case file rather than in a hundred-millisecond authorisation path.
- Medical Coding & Claims AgentsThe 835 remittance hands you a free adversarial label on every claim, and it is biased in exactly one direction — undercoding is never denied, so a single-sided objective has a degenerate optimum the optimiser will find; make the scorecard two-sided against a blind-coded baseline, route on the CARC/RARC pair because CO-16 is a pointer rather than a reason, and treat autonomy as a jurisdiction-keyed routing decision now that Indiana requires human review before submission.
- RFP & Security-Questionnaire AgentsEvery answer is a contractual representation, so a stale claim is a misrepresentation, not a hallucination: an answer library with an owner and a valid_until per claim, matching that scores the difference not the similarity, three tiers with a hard block on future-tense commitments.
- Expense & Travel Audit AgentsThe recoverable money is unreclaimed VAT, duplicates and priced-in policy leakage — not fraud — so scale the deterministic checks to 100% and leave the judgement calls sampled, keep policy in an effective-dated rule engine the model never reads, and score the one job the model owns (reconciling receipt, card feed and claim) against a card-settlement ground truth that arrives free every night.
- Clinical-Trial Matching AgentsThe expensive error is invisible — an eligible patient who was never surfaced — so build for recall, not for the screen-failure rate you can see: decompose free-text criteria into atomic predicates that each carry a time window, emit a per-criterion table with evidence spans and an explicit unknown state instead of a verdict, and gate patient contact behind a named human.
- Pharmacovigilance & Adverse-Event AgentsPutting an agent on a channel is a legal act, not a monitoring decision: the reporting clock starts at first knowledge by anyone acting for the company, so the written channel scope comes before the model. Build around the four elements that make a case valid, emit missing-element follow-up questions instead of a reportable yes/no, constrain coding to the dictionary version in force, and let the agent escalate while only a qualified person may dismiss.
- Drive-Thru & Restaurant Ordering AgentsVoice AI alone gets about 83% of drive-thru orders right against 87% for the standard lane, and about 95% when staff step in — which they do on roughly one order in five — so the product is the handoff, not the recogniser. Track containment, intervention cost, order time and drive-offs rather than accuracy; ground every item in the live menu version and write through the POS; and remember the hard part is the modifier grammar and the lane’s audio, not the transcription.
- Ambient Clinical Documentation AgentsThe headline result moved burnout from 51.9% to 38.8% and did not measure documentation quality at all — so time-to-signature, the metric every dashboard ships, is maximised by a clinician who signs without reading. Design backwards from the signature as an authorship transfer, enforce provenance tiers so a physical-exam finding appears only if it was said aloud, surface omissions because review cannot catch them, keep code assignment out of the scribe, and fund the weekly adjudicated sample that is the only number measuring the note.
- Prior Authorization AgentsStatute has already split this workflow: across eleven states in two sessions an AI may not be the sole basis for a medical-necessity denial, while the same laws expressly permit AI for administrative work and organising clinical information — so build an agent that can approve and escalate but cannot deny, with the denial branch absent rather than flagged off. Since January 2026 payers owe 72 hours expedited and 7 days standard with a specific reason attached, so engineer against the reason code and first-pass completeness rather than an approval-probability model — and read WISeR, where Texas requests ran 62% approved by machine and 84% after human review, as a lesson about the objective rather than the technology.
- Insurance Underwriting AgentsBind rate arrives in seconds and loss experience arrives after the development tail, so any loop closed on the fast signal is an adverse-selection machine — and the regulated artifact is the enumerable factor that sets price, which an LLM judgement can never be. Emit declared variables with a source span and an explicit unknown, leave the rating engine deterministic, run the BIFSG-style outcome test on your own book before an examiner does, and make the decline path a referral rather than a branch.
- Smart-Home & IoT AgentsEvery demo turns on a light and every incident will be about something the agent read: device event history is a presence log for the whole household, attacker-writable at the device-name layer and impossible to un-read once it is in a context window. Classify devices by reversibility rather than category, write policy over runtime-discovered traits with default-deny, express commands as absolute targets verified by observation, cap the history window server-side, and design for the people in the building who never saw the consent screen.
- Consumer Credit AgentsTeams brace for the explainability problem and get caught by the conversation: under Regulation B what the agent asks is restricted, what it says to a hesitant applicant can be unlawful discouragement, and the moment it stops collecting documents starts a thirty-day notice clock nobody wired a timer to. Build intake as a closed question registry with a coded completeness rule, keep scoring in a declared model the agent cannot reach, and measure document-extraction accuracy on the denial population — that tail is what produces wrong principal reasons.
- Payroll & Employment-Tax AgentsPayroll already contains a deterministic authority, so an agent that produces a figure destined for a filing has inserted a probabilistic step into a process whose whole penalty structure assumes a determinate one. Keep it on the input and reading sides of the engine, and note that the primary control is a timer rather than a confidence threshold: US deposit penalties step 2/5/10/15% by lateness alone, so an agent that asks a question and waits has produced a compliance failure. Spend the capability on notice triage, variance and reconciliation, and never let a model infer one of 7,400+ local jurisdictions.
- Scientific-Discovery AgentsTwo benchmarks posted to arXiv in early October 2026 settle the design question: on EurekaBench an agent hit 47.4% predictive accuracy against a human scientist's 48.8% and 29.4% on the insight the task was built around against 69.7% — so build for the metric it already wins and you ship expert-accuracy correlations nobody can publish. Split execution, analysis and mechanism-proposal into separate agents, pre-run the expensive simulations and grade with fixed rules rather than a judge, declare guidance as a logged L0–L3 parameter, and report expert-accepted insights per expert review-hour.