Prior authorization agents.
Statute has already split this workflow in half and handed your agent the side nobody is selling. Across a growing list of US states an AI may not be the reason a patient is told no — medical-necessity denials belong to a licensed clinician — while the same laws explicitly leave administrative work and the organising of clinical information open. So the buildable agent is the one that can say yes, can escalate, and structurally cannot deny; and since January 2026 the payer owes a decision on a clock with a specific reason attached, which means you should be engineering against the reason code, not against an approval-probability model.
The clock is already binding. The API that was supposed to make it survivable is not here yet.
CMS's Interoperability and Prior Authorization final rule (CMS-0057-F) restructured the timing side of this workflow first and the plumbing side second, and the gap between those two dates is the whole reason an agent is worth building right now.
- Since 1 January 2026, impacted payers owe decisions in 72 hours for expedited requests and 7 calendar days for standard ones, and must give a specific reason when they deny. Impacted payers means Medicare Advantage organisations, state Medicaid and CHIP fee-for-service programmes, Medicaid and CHIP managed care, and qualified health plan issuers on the federally-facilitated exchanges.
- Those exchange issuers are carved out of the timeframes — they stay on the older 15-day/72-hour standard — but not out of the denial-reason and public-reporting duties. Your routing table needs that distinction in it, because it is the kind of exception that becomes a compliance finding when it is hard-coded as a single rule.
- Payers must now publish prior-authorization metrics annually — volumes, approval and denial rates, appeal outcomes, average turnaround — with the first report due 31 March 2026 covering calendar year 2025. This is free competitive intelligence about the specific payers you submit to, and almost nobody on the provider side is reading it.
- The four FHIR APIs do not land until 1 January 2027, including the Prior Authorization API that is supposed to let you query documentation requirements, submit, and read status and reason programmatically — for all items and services except drugs.
So for this year you are automating against portals, faxes and phone trees to hit a deadline that is already legally enforced. That is worth doing, and it sets a design constraint: the scraping and form-filling layer has a known expiry date, while the layer that decides what evidence to gather does not. Build them as separate components, and expect to throw one away.
There is a follow-on proposed rule, CMS-0062-P, published in April 2026, extending electronic prior authorization to drugs and shortening timeframes further. As far as we can establish it is still proposed rather than final. Treat drug prior authorization as a separate roadmap item on a separate calendar, not as a phase two of the same build.
The agent's job is the evidence packet, not the prediction.
The tempting architecture is a model that reads the chart and predicts approval, so the practice can triage. It is the wrong objective in a domain where the criteria are, increasingly, published. What you actually want is a much less glamorous question answered exactly: does this chart contain the specific things this payer's rule asks for, and if not, which ones are missing?
The standards stack is built around that question, and it is worth knowing which piece answers what:
- Da Vinci CRD (Coverage Requirements Discovery) fires from the EHR at order or scheduling time and answers whether authorisation is required at all, and whether special documentation rules apply. Catching "no PA needed" here is the cheapest win in the entire workflow.
- Da Vinci DTR (Documentation Templates and Rules) brings the payer's questionnaire and rules into the EHR and captures the required documentation, auto-populating from the chart. This is the step your agent should be strongest at.
- Da Vinci PAS (Prior Authorization Support) submits the request and manages its lifecycle — status, update, cancel — by wrapping the HIPAA X12 278 transaction inside FHIR. CMS has issued enforcement discretion not to enforce the X12 278 requirement where a covered entity uses the FHIR-based Prior Authorization API, which is what makes the FHIR path viable rather than additive.
Framing the agent as a documentation-completeness engine rather than an oracle buys you three things. The output is checkable — a named missing element is either in the chart or it is not. It degrades honestly, because "I could not find a trial of conservative therapy in the record" is useful even when wrong. And it is squarely inside what every AI-in-utilization-review statute leaves permitted, which is the subject of the next step.
Ground every extracted element in the document it came from, with a pointer a human can open. A prior-authorization packet asserting a date of onset the chart does not support is not a quality problem, it is a false statement on a benefits request — the grounding discipline here is non-negotiable, and document parsing quality sets the ceiling on everything above it.
The autonomy line is drawn by statute, and it is drawn at "no".
California's SB 1120, the Physicians Make Decisions Act, has been in effect since 1 January 2025 and is the template the rest of the country has been copying. It requires that a denial, delay or modification of care based on medical necessity be made by a licensed physician or other competent licensed professional; AI used in utilization review may not supplant provider decision-making. Critically for anyone building here, it expressly does not prohibit using AI for administrative tasks or for organising medical information.
The trend has accelerated. Four states passed laws in the 2025 session barring AI as the sole basis for a medical-necessity determination — Texas, Maryland, Nebraska and Arizona among them — and seven more passed such laws in 2026, including Alabama SB 63, Colorado HB 1139, Indiana HB 1271, Utah SB 319 and Washington SB 5395. Federal statute is absent; a December 2025 executive order seeks to preempt state AI laws through litigation and funding conditions, and health-insurance AI rules are not among its named carve-outs, so the durability of these protections is genuinely uncertain. Meanwhile Medicare Advantage rules already predating the AI wave bar making medical-necessity decisions with an algorithm that does not consider individual circumstances, and require a health care professional to review such denials.
Read together, these give you a design rule rather than a compliance checkbox, and it is asymmetric in the same way adverse-event intake is:
- The agent may approve, assemble and escalate. It may never deny. Build the denial path so it is not reachable — no branch, no confidence threshold, no configuration flag that turns it on. A control that exists but is disabled is a control someone will enable during a backlog.
- "Not approved by the agent" is a routing outcome, not a decision. It means a qualified human now owns the case. Name the state that way in your schema, because the word you choose will end up in a screen, and eventually in a deposition.
- Store the reviewer as a fact, not an inference. Which licensed professional reviewed this, when, and what they saw. Timestamps adjacent to a queue are not evidence of review — the audit trail has to record the act.
- Make jurisdiction a configuration with an effective date. Eleven states in two sessions means you will edit this several times a year, and it belongs in data rather than in a conditional.
Watch what the objective is pointed at — WISeR is the published cautionary case.
CMS's own Wasteful and Inappropriate Service Reduction model launched on 1 January 2026 and runs through 2031 in six states — Arizona, New Jersey, Ohio, Oklahoma, Texas and Washington — applying AI-assisted prior authorization to Traditional Medicare for a specific list of services, with private technology vendors paid a share of the Medicare spending reductions they produce.
That last clause is the design lesson, and it has nothing to do with model quality. A system whose vendors are compensated out of reductions has an objective pointed at reductions, and you do not need bad faith or a bad model for that to show up in the numbers. It did:
- In Texas, 62% of WISeR requests were initially approved, rising to 84% after human review, against a roughly 92% national approval rate in Medicare Advantage. The 22-point gap between the machine's first answer and the reviewed answer is the measurement of the objective, not of the technology.
- The oversight record is public and unusually blunt. GAO determined in May 2026 that the model constitutes a rule under the Congressional Review Act, partly because it transfers decision-making authority over Medicare claims to entities using AI that are paid out of spending reductions. A resolution to repeal failed in the Senate on 16 July 2026 by 46–50, leaving the model running.
- Records released in September 2026, obtained through litigation over algorithmic transparency, describe backlogs, testing gaps, widespread delays and reports of patient harm.
If you are building on the payer or vendor side, the transferable rule is that the compensation structure is part of the model specification and will be read that way by regulators. If you are building on the provider side, WISeR is an operational fact: in those six states, for those services, expect a first answer that human review moves substantially, which makes your appeal path a primary workflow rather than an exception handler.
The label is asymmetric, delayed, and mostly wrong in one direction.
Prior authorization looks like it hands you clean supervision: you submitted, you got approved or denied. It does not, and the distortion runs the same way every time.
- A denial is a weak negative. On the AMA's 2025 physician survey, 81.7% of appealed denials were overturned in full or in part. Training or tuning on first-pass outcomes teaches your agent to reproduce a decision that is reversed most of the time it is contested.
- An approval is silent about quality. It tells you the packet cleared, not that it was efficient, not that the service was appropriate, and not that a thinner packet would have cleared too.
- The real label arrives weeks later. The appeal outcome is the strong signal and it matures long after the dashboard has drawn its line. Hold the cohort and evaluate on matured cases — the same discipline as claims and coding, where the delayed-label trap is identical.
Because of that, the metric to optimise is not approval rate. It is first-pass completeness: the share of submissions that never generate a request for additional information. That metric is available immediately, it is not confounded by payer behaviour, and it is the one the agent can actually move. Pair it with a missing-element taxonomy — which specific required element was absent, by payer and service — because that distribution is your roadmap, and it is the direct analogue of a failure taxonomy.
The burden numbers explain why even modest movement here pays. The same AMA survey found physicians and staff handling about 39 prior authorizations per physician per week and spending roughly 13 hours a week on them. Per-transaction industry cost data puts a manual prior authorization at several times the cost of a fully electronic one for the provider, with manual submissions consuming roughly 24 minutes each. The return is in requests that never bounce, not in requests that get approved faster.
Ship it against the 2027 target, because that is the brief already written for you.
Under the industry pledge announced in June 2025 — roughly 60 insurers covering some 257 million Americans — plans committed to a standardised FHIR-based electronic submission framework and to answering at least 80% of electronic prior-authorization approvals in real time by 1 January 2027, when all needed clinical documentation is attached. That conditional clause is your product specification. The value of the agent is making it true on the first submission.
Hold the pledge at arm's length while you build to it. On the AMA's 2026 survey, only 33% of physicians believed the commitments would make a meaningful difference, and only 24% reported that medical-necessity denial reviews were consistently performed by appropriately qualified clinicians. Plan for the standard, not for the goodwill.
- Sequence by reversibility. Start where a wrong answer costs minutes: CRD-style "is authorisation even required" checks, and documentation-gap detection surfaced to staff. Move to auto-submission only for service and payer combinations where your first-pass completeness has been high for months.
- Instrument the clock as a countdown, not a log. The 72-hour and 7-day obligations are the payer's, but your appeal leverage depends on knowing when they lapse. A submission timestamp that nobody is counting down from is an unused right.
- Read the published metrics. Payers now post approval and denial rates and turnaround times annually. Feed them into routing: a payer whose denial rate for a service is an outlier deserves a fuller packet, and the evidence for that judgement is now public.
- Keep the human path staffed as you scale. Automation that increases submission volume increases escalation volume proportionally, and the escalation queue is the one that has a licensed professional at the end of it — see human in the loop for keeping that a control rather than a rubber stamp.
Build the agent so that the only outputs it can produce are "submit this packet", "this packet is missing X", and "a clinician must look at this" — and make the denial branch literally absent from the code rather than gated behind a flag. Then measure first-pass completeness by payer and service, hold your evaluation cohorts until appeals mature, and keep the portal-scraping layer separate from the evidence-assembly layer, because the first one has a 2027 expiry date and the second one is the actual product. Related: healthcare agents for the clinical-safety frame, insurance claims agents for the same workflow seen from the payer's desk, accountability and roles for naming who owns the determination, and model risk management for the validation programme a regulated deployment needs.