Pharmacovigilance & adverse-event intake agents.
Pointing an agent at a channel in drug safety is not a monitoring decision, it is a legal one: the reporting clock starts when the company first becomes aware of a case, and a system that reads a mailbox makes the company aware of everything in it. Build the channel scope before you build the model, tune for recall and abstention rather than precision, and never let the agent be the thing that decides a case is not reportable.
Reading a channel is an act with a deadline attached.
Post-marketing safety reporting runs on a clock that starts at first knowledge. Under the US rules an applicant must submit a report of an adverse experience that is both serious and unexpected no later than fifteen calendar days from the initial receipt of the information, and "receipt" means by anyone acting on the company's behalf — not by the safety department, not by the person who eventually triages it. The EU shape is the same idea with a second tier: fifteen days for serious cases, ninety for non-serious.
Now put an agent on a social channel, a call-centre transcript feed, a patient-support inbox or a distributor's ticket queue. Everything in that channel is arguably now known to the company. You have not merely built a detector; you have expanded the surface over which day zero can be triggered, and you have done it at machine scale on material nobody was previously obliged to read.
So the first artefact is not a prompt. It is a written channel scope, approved by whoever owns the reporting obligation, that says which sources the agent reads, which it does not, and why. That document is the thing an inspector will ask for, and it is the thing that stops a well-meaning integration from quietly doubling your case volume on a Friday.
The corollary that surprises teams: adding a channel is a regulatory change, not a feature flag. Treat a new connector the way you would treat a new study site — reviewed, documented, dated. Third-party tool drift is the failure mode where a vendor quietly widens what a connector returns and your scope document silently stops being true.
The unit of work is a valid case, and validity is four slots.
An individual case safety report becomes reportable when four elements are present: an identifiable reporter, an identifiable patient, a suspect medicinal product, and an adverse reaction. Miss one and you do not have a case — you have a signal that needs follow-up. That structure is the single most useful thing about this domain, because it turns a fuzzy classification problem into a slot-filling problem with a known schema.
Build the agent around the slots, not around a yes/no verdict. For each incoming item it should emit the four elements with an evidence span for each, an explicit missing or unclear state where the text does not support a value, and — this is the part that earns its keep — the specific follow-up question that would complete the case. "Reporter identifiable, patient identifiable, product named as Brand X, reaction described as persistent cough; onset date absent — ask when the cough began" is a work item a case processor can act on in seconds.
What the agent must not do is compress. The verbatim reporter text is a regulated artefact and has to survive intact into the case record alongside anything the model derived from it. Structured output is an addition to the source, never a replacement for it; the same discipline as structured outputs generally, with a much sharper penalty for losing the original.
Tune for recall, and make abstention a first-class answer.
The two error types are not comparable here. A false positive costs a few minutes of a case processor's time. A false negative is an unreported case, which is a compliance finding, a potential inspection observation, and — upstream of all of that — a missed contribution to a signal that might have changed a label. Optimising a threshold on F1 quietly trades the second for the first at a ratio nobody agreed to.
Three consequences for the build:
- Report recall, stratified. A single aggregate hides the cases you most need to see. Break it out by channel, by language, by product, and by whether the reaction was described in lay terms or clinical ones — lay-language recall is usually the weakest and is exactly where consumer reports live.
- Make "I am not sure" a routed outcome, not a low score. An item the agent cannot resolve should land in a human queue with the reason attached, and the rate of that routing is a number you publish. Uncertainty and calibration is what lets you price it.
- Never let the agent close. The agent may promote an item to a human; only a qualified person may decide an item is not a valid case. That asymmetry — machine escalates, human dismisses — is the whole safety property, and it survives every future model upgrade without renegotiation.
Expect the volume conversation. A recall-first intake agent will raise apparent case volume, and somebody will read that as the system being noisy. It is worth saying in advance, in writing, that the increase is the point: you were previously not reading those channels.
Coding is a controlled-vocabulary problem, and that is where models quietly lie.
Reactions get coded to a controlled dictionary — MedDRA in practice — and the coded term is what aggregates into signal detection. This is the step where a language model looks most useful and is most dangerous, because a plausible term is not a valid one, and an invalid term does not error: it simply lands in the wrong bucket and disappears from the analysis it should have joined.
Two rules keep this safe. First, the model never emits a term; it retrieves candidates from the dictionary version in force and proposes a ranked shortlist with the verbatim span that motivated each. Constrain generation to the vocabulary rather than trusting the output — the mechanism is constrained decoding, and the alternative is a hallucinated code that reads correctly to everyone downstream.
Second, treat the dictionary version as part of your system version. Vocabularies are revised on a schedule; a term that was current last release may be demoted this one, and a comparison drawn across the change is not a comparison. Pin it, stamp it into the record, and re-baseline your evaluation when it moves — the same tuple discipline as rollout and versioning.
The evaluation to build first is a coding agreement set: a few hundred verbatim texts with adjudicated terms, scored as top-1 and top-5 agreement against your own coders. Top-5 is the number that tells you whether the shortlist is useful; top-1 is the number that tells you whether anyone will trust it enough to accept without reading.
Seriousness and expectedness set the clock — so keep them out of the model's hands.
Two determinations drive the deadline. Seriousness is criteria-based: death, life-threatening, hospitalisation or its prolongation, persistent or significant disability, congenital anomaly, or a medically important event. Expectedness is a comparison against the reference safety information for that product. Together they decide whether an item is a fifteen-day case or a routine one.
Seriousness is genuinely tractable for an agent, because it is a checklist against text, and the criteria are stable. Build it as explicit per-criterion extraction with evidence, not as a single label. Expectedness is different: it requires reading the current reference document for the right product and the right version, and it is a judgement that belongs to a qualified person. The agent's correct output here is a prepared comparison — reaction term, the relevant section of the reference text, and what differs — not a verdict.
Put the clock itself in the system rather than in a spreadsheet. Day zero is a timestamp: the moment the item entered a channel in scope. Record it at ingestion, propagate it through every downstream state, and alarm on age rather than on queue depth. A queue that is short because everything in it is eleven days old is the exact failure this system exists to prevent, and it looks healthy on a dashboard that counts items — see production feedback signals for choosing the signal that would actually have fired.
Validate it as a regulated system, because that is what it is.
This system participates in a regulated process, which means the questions asked of it are not the ones usually asked of an agent. Four constraints are design decisions, cheap early and expensive later.
- Reproducibility is mandatory. Given the same input and the same configuration, the output must be reconstructable when someone asks about it eighteen months from now. Pin model versions, store prompts with the record, keep the retrieved reference spans, and treat a model swap as a change to a validated system — with impact assessment and re-qualification, not a silent upgrade. Model deprecation and migration is the part of this you do not control and must plan for.
- Audit trails are per decision, not per run. Who or what proposed each element, what evidence supported it, who accepted or overrode it, when. Audit trails and provenance covers the shape.
- Accountability has a named holder. In the EU a qualified person is personally responsible for the pharmacovigilance system; the agent is a component inside their system, and they need to be able to describe what it does. Write that description for them, in their language, before the first release.
- Minimise what leaves the building. Case narratives are identifiable health data with a reporter attached. De-identify before any call that leaves your boundary, keep the re-identification map inside it, and check the retention terms of every endpoint in the path — zero data retention and abuse monitoring is where that gets decided, and it is decided per endpoint, not per account.
Ship the narrowest thing that is useful: one channel already in scope, one product family, intake only. Output the four validity elements with evidence spans, an explicit missing state, a proposed follow-up question and a seriousness checklist — and no verdicts, no coding decisions, no closures. Measure stratified recall against an adjudicated set before you measure anything else, publish the abstention rate next to it, and instrument day-zero age from the first week. The teams that get into trouble here are the ones that started with a classifier that answers "reportable: yes/no", because that is the one output nobody can audit and the one decision the agent is not allowed to make. Related: healthcare agents for the surrounding rails, medical coding & claims agents for the controlled-vocabulary pattern this shares, and content moderation agents for the queue mechanics.