Ambient clinical documentation agents.
Every number that sold ambient scribes measures the clinician, not the note — the headline JAMA Network Open result moved burnout from 51.9% to 38.8% and did not measure documentation quality at all. That gap is the whole design problem: time-to-signature and note fidelity pull in opposite directions, only one of them is instrumented, and the signature is an authorship transfer that makes a human legally responsible for text a model wrote. Build the review step, not the draft.
Pick the KPI before you pick the vendor, because the obvious one is adversarial.
Documentation burden is real and ambient capture genuinely relieves it; clinicians report thirty minutes to two hours a day back, and roughly a third of providers now have access to something in this category. None of that is in dispute. What is in dispute is what you should be optimising once it is deployed.
"Minutes saved per encounter" and "time to signature" are the metrics every dashboard ships with, and they are maximised by a clinician who signs without reading. The same evidence base is already ambiguous on the longer horizon — a University of Pittsburgh follow-up found burnout rising again across later time points — which is what you would expect if the first-month gain comes partly from reading less and the later cost comes from correcting notes downstream.
- Treat time saved as a constraint ("must not increase documentation time"), never as the objective. Objectives get gamed; constraints get satisfied.
- Make the objective a fidelity measure that only exists if you build it: what fraction of signed notes contain a clinically material error or omission, sampled and adjudicated. Nobody ships this by default.
- Expect the two to trade. A programme that improves both in the same quarter is usually measuring signature latency and calling it quality.
Design backwards from the signature, which is an authorship transfer.
The clinician who signs the note is its author for every purpose that matters afterwards — clinical hand-off, billing substantiation, malpractice discovery, licensure. The model has no standing in any of those rooms. This is not a governance footnote; it dictates the interface.
- Make the note reviewable in less time than it takes to write. That is the actual product requirement, and it is a UI requirement, not a model one. Diff-style presentation, each assertion linked back to the moment in the encounter that produced it, and the unsupported lines visually separated — see review queues for agent output.
- Never pre-fill the signature path. One-click sign-all across a day's encounters converts the review step into a formality, and it is the single most requested feature. Decline it, or rate-limit it to notes below a change threshold.
- Keep the draft attributable. The record should carry which model version and which prompt produced the draft, and what the clinician changed before signing — the ordinary audit trail argument, with the unusual property that the artefact under audit is a legal document about a person.
A useful test for any proposed feature: would you be comfortable if it were described accurately in a deposition? "The system highlighted three statements it could not support and the physician reviewed each" survives that. "The system produced a complete note and the physician signed it in four seconds" does not.
You have created a second record, and nobody has decided what happens to it.
Ambient capture means a recording of the encounter now exists. Whatever the note becomes, the audio and the transcript are separate artefacts with their own consent, retention and discovery consequences — and teams routinely ship without a decision on any of them.
- Consent is per jurisdiction and per encounter, not per deployment. All-party-consent states change the script; so does a third party in the room. Build the consent capture into the encounter start, record which version of the disclosure was read, and make declining leave a working fallback — a patient who says no cannot be a patient who waits. The mechanics mirror recording consent & redaction.
- Decide the audio retention window before launch, not after the first subpoena. Keeping raw audio indefinitely gives you a debugging corpus and gives opposing counsel a recording of everything said in the room, including the part that is not in the note. Most programmes should delete audio on signature and keep the transcript only as long as the quality sampling in STEP 6 requires.
- Traces are PHI. The prompt, the transcript and the model output all contain protected health information, which rules out the default observability posture — see PII redaction in agent traces and retention & legal hold.
The dangerous error is the omission, not the invention.
Fabricated sentences are the failure everyone anticipates and the easiest to catch: a clinician reading their own note notices a symptom the patient never reported. The errors that survive review are quieter. Informatics researchers have named the characteristic one hallucination by simplification — the summary flattens a hedged, conditional, or negated statement into a clean declarative one, and the clean version reads like a better note.
- Enforce a provenance tier on every assertion. Tier one: stated in the encounter, with a timestamp. Tier two: pulled from a structured EHR field, with the field named. Tier three: inferred — which the note may not contain at all. Grounding is the whole job here; see hallucination & grounding.
- Ban asserted negatives that were never uttered. "Lungs clear to auscultation" when no examination was narrated is the canonical dangerous output, because it is both plausible and billable. A physical exam finding may only appear if it was said aloud.
- Preserve hedges and negations verbatim. "Denies chest pain" and "no chest pain reported" are different claims about who said what. Where the model would smooth, it should quote.
- Surface omissions explicitly. The note should end with what the model heard and did not place — the allergy mentioned in passing, the medication the patient said they stopped. An omission the reviewer is shown is a caught error; an omission is otherwise invisible by construction.
Do not let the thing that writes the note also decide what it is worth.
The note is billing substantiation. A richer note supports a higher evaluation-and-management level, so any system that is rewarded — even implicitly, even by a vendor's ROI dashboard — for revenue per encounter has a one-directional incentive to write more. Left alone, an optimiser finds that direction, and the result is documentation that is defensible sentence by sentence and indefensible as a pattern.
- Separate the roles. The scribe produces the clinical narrative. Code assignment is a downstream step with its own controls and its own two-sided scorecard, as medical coding & claims agents sets out — undercoding is never denied, so a single-sided objective has a degenerate optimum.
- Watch the distribution, not the note. Audit exposure shows up as a shift in your E/M level mix after go-live, by clinician and by specialty. Chart-by-chart review will not find it; a monthly distribution comparison against the pre-deployment baseline will.
- Refuse template inflation. A per-clinician style template that always emits a ten-point review of systems is copy-forward with a new name, and it defeats the provenance tiers in STEP 4 by construction.
Measure what survives the signature, and keep some encounters off the system.
Four numbers and one exclusion list are enough to run this responsibly.
- Edit rate at signature, by clinician. Not as a quality score — as an attention detector. A clinician whose edit rate falls to near zero in week three has stopped reading, and that is an intervention about training, not about the model.
- Material-error rate on a sampled, adjudicated set. Draw signed notes weekly, have a second clinician compare each against the transcript, and classify: fabrication, omission, distortion. This is the only number in the programme that measures the note.
- Downstream correction rate. Addenda and amendments filed after signature, which is where the cost of a fast signature reappears weeks later.
- Patient-visible outcomes. Notes are released to patients; complaints about note content are a free, biased, and very fast signal.
- An explicit exclusion list. Encounter types where ambient capture is off by default — behavioural and mental health, encounters involving suspected abuse, anything where the recording itself changes what the patient will say. The list is a clinical decision, made once, in writing, and enforced in code rather than by policy.
Ship it as a drafting tool with an aggressive uncertainty display and no bulk-sign, run it for a quarter on one specialty, and fund the weekly adjudicated sample from day one — it is the only instrument that will tell you whether the programme is working, and it is the first line cut from every business case. If you can only build one thing beyond the draft, build the omission list: it turns the failure mode that review cannot catch into one that it can.
Related: healthcare agents for the wider clinical surface, human in the loop for where the gate belongs, and adapting a playbook for deriving the next vertical.