Clinical-trial matching agents.
The expensive error in trial matching is the one nobody ever sees: a patient who was eligible, was never surfaced, and quietly went onto standard of care. A false positive costs a screening visit and shows up in your screen-failure rate within a fortnight; a false negative costs a person their option and shows up nowhere, which is why a system tuned on precision will look excellent and do the opposite of its job. Build for recall, output per-criterion evidence rather than a verdict, and make "unknown" a first-class answer.
Decide who the user is, because it settles the error budget.
There are two products here and they are not variations of each other. One serves a research coordinator with a protocol and a patient list, asking which of these people might qualify. The other serves a clinician or a patient asking which of some 400,000 registered studies might fit this person. The first is a filter over a known cohort; the second is a search over a registry. Build one.
- Coordinator-facing is where to start. The user is an expert, the protocol is fixed, the output goes into a workflow that already exists (pre-screening), and there is a natural ground truth arriving weeks later in the form of who actually enrolled. Every one of those properties is missing from the patient-facing version.
- The agent pre-screens; it never screens. Eligibility is a clinical determination made by a qualified investigator against the protocol. What you are building produces a ranked, evidenced shortlist for a human to work through — and saying so precisely is not a disclaimer, it is the system boundary that keeps this out of regulated device territory in most jurisdictions.
- Never let the agent contact a patient. Outreach is where an unreviewed match becomes a real-world harm: a false hope, an unsolicited disclosure of a diagnosis to a household, a protocol violation. The gate before contact is a named human, always. See human in the loop for where the gate sits.
The asymmetry is the design input. A false positive is absorbed by a process that already expects them — screen-failure rates of around 30% are routine and have exceeded 70% in some Alzheimer's programmes. A false negative is absorbed by nobody, is never measured, and is the outcome the whole enterprise exists to prevent. Tune accordingly, and expect that to feel wrong to anyone reading a precision number.
The product is criteria parsing, not matching.
Matching is trivial once both sides are structured. Neither side is. Eligibility criteria are free prose written for a human reader, and they are long and getting longer — across NCI-affiliated trials they run from a dozen words to around 750, and the trials in the wordiest decile fail to accrue at roughly 29% versus 12% in the shortest. That prose is the thing your system has to turn into something checkable.
- Decompose into atomic predicates, one criterion per row. "ECOG 0–1, adequate renal function, no prior anti-PD-1" is three checks with three different data sources and three different failure modes. A single embedding of the whole criteria block is the most common architecture here and it is the reason these systems cannot explain themselves.
- Every predicate carries a time window. "No systemic therapy within 4 weeks" is not a fact about the patient, it is a fact about a date range relative to a screening date that has not happened yet. Criteria that parse fine and silently drop the window are the single largest source of confident wrong matches.
- Separate the criteria that a computer can decide from those it cannot. Age, lab values, ICD codes and medication history are decidable. "Life expectancy greater than 12 weeks", "clinically significant cardiac disease", "able to comply with the protocol" are investigator judgements that no amount of retrieval turns into a boolean. Marking them requires clinician review is the correct output, not a weakness.
- Treat the registry as a drifting upstream. Protocols amend, sites open and close, recruitment status changes without the record changing. Snapshot the registry, version it, and record which snapshot produced a given match — see third-party tool drift for the general shape of this problem.
Never emit "eligible". Emit a criterion table with evidence.
A verdict is unfalsifiable and unreviewable, and a coordinator who cannot see why cannot do anything with it faster than reading the chart themselves. The output that saves real time is a row per criterion with three columns: the predicate, a state, and a pointer into the record.
criterion state evidence -------------------------------------------------------------- age >= 18 met DOB 1974-03-02 ECOG 0-1 met note 2026-08-14 "ECOG 1" eGFR >= 60 met lab 2026-08-29: 72 no prior anti-PD-1 unknown no oncology history 2019-2023 LVEF >= 50% not assessed no echo on file life expectancy > 12 weeks clinician judgement, not derivable
- Three-valued logic, minimum. Met, not met, unknown. Collapsing unknown into not-met is how a recall-critical system silently becomes a precision-optimised one — most patients are unknown on most criteria, because charts are incomplete.
- Unknown is the most valuable cell in the table. It is a work order: get an echo, call the outside hospital for records, ask the patient one question. A system whose output is "12 patients, each blocked on one retrievable fact" is worth more than one that returns three confident matches.
- Every met/not-met needs a span, not a summary. A pointer to the specific lab row or the sentence in the note. Without it the coordinator re-derives the whole thing, and you have added a review step rather than removed one. This is the same discipline as grounding, applied where the cost of an ungrounded claim is a screening visit.
- Rank by how few unknowns remain, not by a similarity score. The useful ordering is the coordinator's actual question — which of these is closest to a decision with the least work?
The record is worse than your pipeline assumes.
Most of the accuracy in production is won or lost on the patient side, not the protocol side, and it is won in unglamorous places.
- The decisive facts live in notes, not in fields. Stage, prior lines of therapy, performance status and the reason a drug was stopped are narrative. Structured problem lists are stale and are maintained for billing, which means they over-report resolved conditions and under-report active ones.
- Normalise units and reference ranges before comparing anything. Creatinine in mg/dL versus µmol/L, haemoglobin in g/dL versus g/L, assay-specific ranges for the same analyte. A unit error in a numeric criterion produces a confident, specific, wrong answer — the worst available failure shape.
- Never let the model infer a diagnosis the chart does not state. Reasoning from a medication to a condition is a plausible inference and an unacceptable one here, because it manufactures evidence for a criterion. If it is not written, the state is unknown.
- Deduplicate across sources before matching, not after. The same lab arriving from the EHR and from an outside PDF with different dates will silently satisfy or violate a time window depending on which one the retriever happened to surface.
- Handle the negations, because prose is full of them. "No evidence of brain metastases" and "brain metastases" differ by two words and invert an exclusion criterion. This is a known weak point of retrieval-first designs and a good reason to keep a deterministic check on the exclusion list.
Evaluate on recall against an adjudicated set, and report screen-failure separately.
Accuracy is meaningless here — the negative class is almost everything, so a system that returns nothing scores above 99%. Two numbers matter, and they are measured against different ground truths.
- Recall against a human-adjudicated gold set. Take a few hundred real patient–trial pairs, have two coordinators independently determine eligibility with the chart in front of them, adjudicate disagreements, and freeze it. This is expensive and it is the only defensible number you will have; budget it as a build cost, and see annotation and labeling ops for how to run it.
- Per-criterion accuracy, not per-patient. A patient-level verdict hides which predicate class is broken. Break results out by criterion type — demographic, lab, medication history, temporal, judgement — because the temporal ones will be your worst category and you will not see that in an aggregate.
- Screen-failure rate is the production signal, and it lags. It tells you about precision, arrives weeks later, and is confounded by everything else in the funnel. Track it, and never let it become the optimisation target: the cheapest way to improve it is to stop surfacing borderline patients, which is the failure you are trying to avoid.
- Measure the unknown rate as a first-class metric. Rising unknowns mean a data-access problem, not a model problem, and the two have completely different fixes.
- Audit the misses, not the hits. Once a quarter, take patients who enrolled on a trial and ask whether your system had surfaced them. That retrospective is the only view you get of the false negatives, and it is worth more than any online metric.
The rails are regulatory, and they are architecture.
This system touches identifiable health data, a research protocol and a person's treatment options at once. The constraints are not paperwork attached at the end; each one is a design decision made early or an expensive retrofit later.
- Match reproducibly or not at all. Given the same patient snapshot and the same registry snapshot, the output must be reconstructable months later when someone asks why a patient was or was not surfaced. Pin model versions, store the retrieved spans with the decision, and keep the criteria parse — see audit trails.
- Pre-screening under a waiver is not recruitment. Reviewing records to identify potentially eligible patients and contacting them are governed differently, and the boundary is usually the point of contact. Build the system so that boundary is a literal gate in the workflow with a recorded approver, not a policy in a document.
- Minimise what leaves the building. Criteria parsing needs the protocol, not the patient. Patient-side reasoning can often run against a de-identified projection, and where it cannot, it should run inside the health system's boundary. Data residency is the surrounding question.
- Watch for equity failures in the miss set. Recall that varies by language, documentation density or site will systematically under-surface exactly the populations trials are already criticised for under-enrolling. Stratify your recall number; a single aggregate hides this completely.
Ship the narrowest version first: one disease area, one site's protocols, coordinator-facing, output a criterion table with evidence spans and an explicit unknown state, and contact nobody. Measure recall against two hundred adjudicated pairs before you measure anything else, and put the unknown rate on the same dashboard. Most teams in this space build a similarity search over criteria text, demo it well, and discover in month six that it cannot say why — the evidence table is what makes the difference, and it is cheaper to build first than to retrofit. Related: healthcare agents for the surrounding rails, uncertainty and calibration for pricing abstention, and research agents for the retrieve-then-verify pattern this shares.