Underwriting agents: assemble the evidence, never set the price.
Underwriting is the one agent deployment where the scoreboard lies for two years. Bind rate arrives in seconds, loss experience arrives after the policy period develops, and anything you optimise on the fast signal is an adverse-selection machine that looks like a triumph until the reserves move. The second constraint is regulatory and just as structural: a regulator does not audit your model, it audits the factors that set eligibility and price — which have to be enumerable, filed and testable. Put the model in the evidence pipeline, keep it out of the rating path, and you get the throughput without making an unfileable variable the centre of your book.
Name the regulated artifact before you name the use case.
In claims, the regulated act is the determination. In underwriting it is narrower and stranger: the regulated artifact is the set of rules and factors by which a risk is accepted, declined or priced — in most US personal lines, filed with the state and in some cases approved before use. The narrative a human underwriter writes in the file is documentation. The factor that moved the premium is the regulated object.
That distinction decides your architecture, because an LLM's output is not enumerable. If a model's judgement becomes an input to price or eligibility, you have created a factor that cannot be listed in a filing, cannot be explained to a market-conduct examiner clause by clause, and cannot be tested in isolation. The problem is not that the model might be inaccurate. The problem is that a filed rating plan has no place to put it.
- Risk selection and pricing are regulated. Eligibility rules, tier assignment, rating variables, surcharge and credit rules, declination reasons.
- Evidence handling mostly is not. Reading the application, pulling the loss runs, ordering the inspection, normalising a schedule of vehicles, chasing what is missing, drafting the file note.
- The line between them is one question. Does this output become an input to eligibility or price? If yes, it is a factor and it is in scope. If no, it is operations.
Commercial and specialty lines are less rate-regulated than personal lines and the temptation to put the model on the decision is correspondingly greater. The filing constraint relaxes; the anti-discrimination and contestability constraints do not, and neither does the two-year feedback lag in STEP 6. Different paperwork, same architecture.
The throughput is in submission handling, and it is enormous.
The underwriting day is not spent judging risk. It is spent opening submissions in eleven formats, reconciling a broker's spreadsheet against the carrier's schema, ordering and waiting for third-party reports, and deciding which of forty submissions to work at all. That work is high-volume, has clean ground truth, and carries almost no regulatory weight.
- Intake normalisation. ACORD forms, broker emails, loss runs, MVRs, inspection reports, financials, SOVs — into a structured submission record with a source span per field. This is document parsing plus structured output, and it is where most of the hours are.
- Completeness and clearance. What is missing for this class and state, is this submission already in the house from another broker, does the named insured match an existing account.
- Triage by appetite, not by risk. "Does this submission match our written appetite, and is it worth quoting" is a routing decision, not a pricing decision. Getting a quote out first is worth real money in commercial lines, and the agent can move that number without touching a rate.
- Referral packaging. When a submission must go to a human, the agent's deliverable is the exhibit: every declared variable with its source document, its page, and an explicit
unknownwhere the evidence is absent.
Notice what all four have in common: a checkable answer. The document either parsed correctly or it did not; the submission either matched the appetite rules or it did not. That is the same reason claims agents earn their keep on completeness rather than adjudication.
Declared variables: the interface that keeps the model out of the rating path.
The whole architecture reduces to one contract. The agent may only emit declared variables — a fixed, versioned list of fields, each with a type, an allowed range, a source citation and an explicit unknown state. A deterministic rating and eligibility engine consumes those variables. The model never produces a score, a tier, a premium, or a recommendation phrased as one.
// Allowed: a declared variable with provenance and an unknown state.
{ "roof_year": { "value": 2011, "source": "inspection.pdf#p3", "confidence": "high" },
"prior_losses_5y": { "value": 2, "source": "lossrun_2021_2026.pdf#p1" },
"occupancy": { "value": "owner_occupied", "source": "application.pdf#p1" },
"wildfire_zone": { "value": null, "status": "unknown", "reason": "no report ordered" } }
// Not allowed: a judgement that becomes an input to price or eligibility.
{ "risk_quality": "below average", "suggested_tier": 4, "recommend": "decline" }
Three properties fall out of that contract, and each is worth the discipline on its own:
- Every priced factor is enumerable. You can list them in a filing, show them to an examiner, and test each one independently.
- Unknown stays unknown. The single most damaging failure in this domain is a model that fills a missing field from the plausible distribution — the ordinary grounding failure, except the output is a rating variable that goes on to price a real policy. An explicit unknown routes to a human or to an ordered report; a quietly imputed value does not.
- Attribution survives. When a disparity shows up in STEP 4's testing, you can trace it to a variable and a data source. If the model's judgement were itself the factor, you would have a number you cannot decompose and a finding you cannot remediate.
You will be tested on outcomes, not on inputs.
This is the part teams consistently get wrong, because it inverts the usual compliance instinct. Excluding protected characteristics from your inputs does not establish that you are clean. The tests being run against carriers measure results, with the protected characteristic imputed from data you did use.
- Colorado. SB21-169 governs the use of external consumer data and information sources and the algorithms and predictive models built on them. The Division's quantitative testing framework estimates race and ethnicity using BIFSG — a RAND methodology applied to name and geolocation — and asks whether outcomes differ unfairly. It began with life insurance and was extended in October 2025 to private passenger auto and health benefit plans, with full compliance required from affected carriers by 1 July 2026.
- The NAIC model bulletin. Adopted December 2023; by July 2026, 25 states had formally adopted it with around eight more in progress. It requires a written AI systems programme — inventory, governance, risk controls, documentation — producible on examination. The NAIC's AI Systems Evaluation Tool ran as a multistate pilot across twelve states from spring to September 2026, which tells you the examination is becoming a standardised procedure rather than an ad-hoc request.
- Consumer reports. Where a third-party report contributes to an adverse decision, federal adverse-action machinery applies on top of all of the above, with its own notice content and timing.
Engineer for that directly. Run the proxy test yourself, on your own book, on a schedule — before an examiner does, and before a new data source silently becomes a proxy. Treat a disparity the way you treat a failing test: attribute it to a declared variable, quantify that variable's contribution, and decide whether it survives. Model risk management is the governance wrapper; the testing itself belongs in your pipeline, not in an annual report.
A declination is a contestable act; build the reason, do not write it.
Declines and non-renewals generate letters, and those letters are read by regulators, agents and sometimes attorneys. A generated explanation is exactly the wrong artifact here: fluent prose that approximates the reason is a misstatement risk, and two policyholders with identical facts must not receive differently-worded reasons.
- Derive the reason from the rule that tripped. The eligibility engine knows which rule fired and on which declared variable. The notice text is a template keyed to that rule, not a paraphrase of the file.
- Attach the evidence. Source document, page, value. When an applicant contests — and contestability is a right in more jurisdictions each year — the correction has to be possible at the level of a fact, not of a verdict.
- Log the whole path. Which reports were ordered, which variables were populated, by what version of which rule set, at what time. Reconstructing a decision eighteen months later is the ordinary case, not the exception — see audit trails.
- Keep the agent able to refer but not to decline. Straight-through acceptance inside written appetite is low-risk and high-value; straight-through declination concentrates every regulatory and reputational risk in the branch with the least oversight. Build the decline path as a referral, and you have removed a category of incident rather than flagged it off.
Measure what you can see now, because the truth arrives in two years.
Here is the trap that makes underwriting different from every other domain in this section. The outcome you care about — whether the risks you bound were priced correctly — emerges over the policy period and the development tail behind it. The signals available this quarter are quote volume, turnaround, hit ratio and bound premium, every one of which improves when you write business you should have declined.
Any optimisation loop closed on those fast signals selects for the risks your competitors turned down. That is not a modelling error; it is the classic adverse-selection dynamic, arriving faster because the agent quotes faster. So:
- Score the process, not the outcome, in the short run. Field-level extraction accuracy against an adjudicated sample, unknown-rate, referral precision, quote turnaround, rework. These are measurable this week and they are what the agent actually controls.
- Watch the mix, not just the volume. Segment the bound book by the variables the agent touched most and compare its distribution against the prior period. A shift in mix is the earliest visible sign of selection drift, and it shows up long before loss ratio does.
- Keep a referral holdout. Route a small random share of straight-through-eligible submissions to a human anyway, and compare. Without it, you cannot tell whether the agent is accelerating good decisions or manufacturing agreement — the same argument fraud and disputes agents make for an approve-anyway holdout.
- Do not let a dashboard metric become the objective. Hit ratio is a diagnostic. The moment it becomes a target for the system, it is a target for the failure mode, and production feedback signals has the general version of that caution.
Start with intake normalisation on one line of business, with declared variables and an explicit unknown state, and leave the rating engine untouched for two quarters. You will get most of the cycle-time benefit, you will build the provenance trail the proxy testing in STEP 4 needs anyway, and you will not have bet the book on a signal that has not arrived yet. If someone proposes letting the model set a tier, ask them to write the filing language for that factor — the request usually answers itself. Related: KYC and onboarding agents for the intake pattern in a neighbouring regime, and the regulatory landscape for how these obligations fit together.