Insurance & liability for agent actions: who pays was decided months before the incident, by three documents.
When your agent books the wrong shipment, quotes a price that does not exist, or deletes a customer's records, who absorbs the loss is not settled by whose fault it was. It is settled by your model vendor's liability cap, the cap and carve-outs in your own customer contract, and whether your policy affirms or excludes AI — all written long before the run. Fault is a distant fourth, and the artefact that decides whether you can recover under any of the three is the trace. An incident you cannot reconstruct is an incident you pay for. (This is engineering guidance for a legal conversation, not legal advice.)
Build the three-document table for your highest-autonomy agent. It takes an afternoon.
Almost no team can answer "if this agent causes a $400,000 loss next month, who pays?" — not because the answer is hard, but because nobody has put the three governing documents on one page. Do it for one agent, the one allowed to act with the least supervision.
- Upstream: the model and tool vendors' terms. Look for the liability cap (commonly the fees you paid in the preceding twelve months), the disclaimer on output accuracy, and the scope of any indemnity. Vendor indemnities in this market typically address intellectual-property claims about generated output — they are not a promise to cover your customer's business loss. Do the arithmetic once: a tool at $2,000 a month caps at roughly $24,000, which is not a meaningful contribution to a six-figure incident. The vendor-side discipline is third-party model & vendor risk.
- Downstream: your own customer contract. Your cap, the carve-outs that survive it (usually gross negligence, wilful misconduct, data breach, sometimes confidentiality), and whether anything in the agreement describes what the agent is permitted to do on the customer's behalf. If your contract promises an outcome the agent delivers autonomously, you have written a warranty.
- Sideways: your insurance. Technology errors-and-omissions and cyber are the two policies that would ordinarily respond. The only question that matters is whether AI-caused loss is affirmatively granted, silently ambiguous, or expressly excluded.
The output is a one-page table with a number in each row. It will usually show that the upstream row is close to zero, which is the whole point of doing it — the exposure you thought you had shared is sitting with you.
Do the exercise per agent, not per company. Autonomy is what drives exposure, and it varies enormously across a portfolio: a summarisation agent with read-only access and a procurement agent that can commit spend belong in different risk classes even though they are the same model behind the same gateway. Your agent inventory is where that per-agent view has to live, and if the inventory does not record what each agent can do unattended, it will not answer an underwriter's questionnaire either.
"Silent AI" is ending, and the default direction is not in your favour.
Most policies written before 2025 neither mention AI nor exclude it. They respond to professional errors, privacy breaches and cyber events without saying whether it matters that a model produced the error. That silence has been read optimistically by buyers for two years, and it is now being closed on both sides.
- Standardised exclusions exist and are in circulation. ISO published optional commercial general liability endorsements addressing generative AI carrying January 2026 edition dates — CG 40 47, which excludes bodily injury, property damage and personal and advertising injury arising out of generative AI, and CG 40 48, which excludes the personal-and-advertising-injury coverage only, with a companion form applying the exclusion to products and completed operations. Whatever your broker says about market practice, the forms exist and attaching them is now a one-line underwriting decision.
- Technology E&O and cyber are being revised the same way. The direction of travel is that AI risk gets expressly allocated — sometimes granted, sometimes removed — rather than left ambiguous. Either outcome is better than silence; the failure mode is finding out which one you got after a claim.
- The dangerous moment is renewal, not the incident. An exclusion added at renewal changes your risk position with no operational signal at all: nothing in your monitoring will fire, and nobody in engineering will hear about it.
The operational move is small and belongs on someone's calendar: ninety days before renewal, ask your broker in writing for every clause in every policy that mentions artificial intelligence, machine learning, automated decision-making or generative output, and put the answer next to the three-document table. Renewal is a change to your risk posture and deserves the same treatment as a model migration — a dated event with an owner.
Affirmative AI cover exists, is small, and is conditional.
A specialist market has formed since 2025 and it is worth knowing the shape of it, with the caveat that every figure here is a snapshot of a fast-moving line of business — verify current terms before relying on any of them.
- Standalone AI liability. Armilla, a Lloyd's coverholder, launched affirmative AI liability cover underwritten by Chaucer and others, addressing losses from underperformance — hallucinated output, model drift, deviation from expected behaviour — with limits that had expanded to the $25 million range per organisation by early 2026.
- Coverage tied to certification. AIUC underwrites against its own audit standard, in which an AI system is assessed across security, safety, reliability, privacy and accountability before cover is written. The structure matters more than the brand: the assessment is a condition of the policy.
- Carrier endorsements and platform programmes. Established carriers including HSB (part of Munich Re) and Counterpart have written affirmative AI endorsements, and Google Cloud's risk-protection programme pairs cloud customers with carriers including Beazley, Chubb and Munich Re.
Two things to take from this. First, "no one will insure agents" stopped being true; if a broker tells you otherwise, they have not looked. Second, the limits available are modest relative to what an agent with payment authority can do in an afternoon, so cover is a backstop for a bounded exposure rather than a substitute for bounding it. The bounding is done in kill switches, spend ceilings and autonomy levels, and it is cheaper than the premium.
Underwriting now reads your evals, which makes them a financial artefact.
The quiet development in this market is that affirmative cover is increasingly conditioned on a governance assessment, and the assessment asks for things your engineering organisation already has — or embarrassingly does not. The questionnaire converges on five things:
- What can the agent do without a human? The single largest rating factor, and the one you can change fastest. Narrowing unattended authority lowers the premium for the same reason it lowers the risk, which is a rare alignment between a control and its cost.
- Can you stop it? A documented, tested kill switch with a named owner and a measured time-to-stop. "We would revoke the API key" is not a control until someone has done it under time pressure — see kill switches.
- What do you measure, and against what? A maintained eval set with a version history, run before releases, is the evidence that you knew the system's failure rate rather than hoping about it. Eval-set maintenance is what turns that from a folder into a record.
- What happened before? Incident history, with root causes and what changed afterwards. An organisation with three well-documented incidents and three fixes underwrites better than one with none recorded, which reads as an absence of instrumentation rather than an absence of incidents.
- Can you reconstruct a decision? Which brings us to the next step, because this is also the question that decides whether you get paid.
The reframing worth internalising: the eval set, the incident log and the audit trail were built as engineering hygiene, and they have quietly become inputs to a price. If you need a business case for the deployment safety checklist beyond "it is the right thing to do", this is it — the same artefacts now sit on a quote.
The trace is what converts a claim into a payment.
Every one of the three documents in step 1 requires you to prove something under pressure, and all three proofs are the same artefact seen from different angles.
- Against the vendor: that a specific component behaved outside its specification on a specific request. Without the model version, the exact prompt sent, the parameters and the response, you have an assertion and they have a cap.
- Against your customer: that you met your standard of care — the controls were in place, the human approval that policy required was obtained, and the agent operated inside the authority the contract described. Accountability & ownership is where the named-approver requirement comes from; the trace is where you demonstrate it happened.
- For the insurer: a proof of loss with a timeline. When the behaviour started, how many transactions it touched, when it was detected and stopped, and what the exposure actually is. An estimate produced from memory a month later is a weak claim and an expensive one.
What this means in practice is that decision receipts — model version, prompt, tool calls, policy decisions, approvals, all joined by a run ID — stop being an observability nicety and become the instrument that decides which of the three caps applies to you. The audit trail page covers what to capture and how to make it tamper-evident.
The same record cuts the other way, and pretending otherwise is how teams get advised to log less. A complete trace can also show that you ignored a warning, disabled a guardrail, or ran an agent past a limit you had set — which is exactly the evidence a plaintiff wants. The correct response is not to record less; it is to decide retention deliberately, with counsel, before an incident makes the decision look like spoliation. Retention & legal hold is the page for that conversation, and the answer is a policy applied uniformly rather than a judgement made per incident.
Write the ceiling into the contract you can actually point at.
The last move is the one that pays off in every direction at once: state the boundary of the agent's authority somewhere legally visible, and keep the system inside it.
- Describe the authority in the customer agreement. Which actions the agent may take unattended, which require a named human approval, and what the per-transaction and per-period limits are. A described boundary converts an open-ended "the AI did something" claim into a question about whether a specific limit was breached, which is a far better argument to be having.
- Make the limits technical, not aspirational. A spend ceiling enforced by the tool layer, an allowlist of counterparties, a per-run cap — enforcement outside the model, per policy enforcement, because a limit that lives in a prompt is a limit that a persuasive input removes.
- Say who the principal is. Disclosure that a customer is dealing with an autonomous system, and on whose behalf it acts, is increasingly a regulatory expectation as well as a good defensive fact — see the regulatory landscape.
- Connect the incident path. Notification duties to your insurer, to your customer, and in some regimes to a regulator have different clocks, and they start before you know the extent. Put all three on one runbook alongside serious incident reporting and incident response, because the first hour is when the clocks are missed.
Do these four things this quarter, in this order: build the three-document table for your highest-autonomy agent and write the actual numbers in it; ask your broker in writing for every AI-related clause in every policy, ninety days before renewal; make the decision receipt a product requirement for any agent that can move money, change records or talk to a customer; and write the agent's authority ceiling into the customer agreement so there is a limit to argue about. Insurance does not reduce your exposure — it converts an unbounded one into a bounded one you have priced. The reduction comes from the ceiling, and the recovery comes from the trace. Buy the policy after you have both, not instead of them.
Related: unit economics for where the premium lands in the cost of a run, human-in-the-loop for placing the approval that your standard of care depends on, and audit trails for the record everything above rests on.