Impact Assessments for Agent Deployments

9 min read

C25
Operation · Governance & Compliance

Impact assessments for agent deployments.

An impact assessment is a dated description of a system, and an agent is the worst possible subject for one: its behaviour is set by a model version, a system prompt, a tool list and a retrieval corpus that each change on different schedules, mostly without a release. So the document you file in March describes a system that stopped existing in May, and the honest audit answer — "that assessment is stale" — is the one answer you cannot give. The fix is not more detail. It is to write an assessment that can be falsified by a monitor: every conclusion names the measurement that supports it, the threshold at which it no longer holds, and who gets told. Then the document stops being a snapshot and starts being a contract with your telemetry.

STEP 1

Three documents, and the one your vendor does not owe you.

Three assessments get conflated, and the distinction decides who holds the pen.

  • The DPIA — a data-protection impact assessment under GDPR Article 35, owed by the controller when processing is likely to result in a high risk to data subjects. Long-established, and most organisations have a template and a reviewer.
  • The FRIA — a fundamental-rights impact assessment under EU AI Act Article 27, owed by deployers of certain high-risk systems: bodies governed by public law, private entities providing public services, and deployers using high-risk AI for creditworthiness assessment or for risk assessment and pricing in life and health insurance. Where the DPIA already covers an obligation, the FRIA complements it rather than replacing it, and the deployer notifies the market surveillance authority of the results.
  • The internal AI assessment — whatever your own risk function calls it. Not legally mandated, usually the only one of the three that actually gets read by the team building the thing.

Read the second bullet again, because it reverses the instinct most procurement processes encode. A FRIA is a deployer obligation. The lab that trained the model does not owe it to you, the platform vendor does not owe it to you, and no amount of vendor documentation discharges it — the assessment is about your processes, your affected population and your oversight arrangements, none of which the vendor can see. If your compliance plan for an agent deployment is a folder of supplier attestations, you have collected evidence for a document nobody has written.

This page is about the assessment as an operational artefact. For which tier your system falls into and what else the Act asks of you at each one, start from the EU AI Act, for agents; for the control-framework mapping, the NIST AI RMF for agents. Nothing here is legal advice, and the one-sentence summaries above are not a substitute for reading Article 27.

STEP 2

The deadline moved. The obligation did not.

The Digital Omnibus on AI was published in the Official Journal on 24 July 2026 and entered into force on 27 July 2026. It deferred compliance for standalone Annex III high-risk systems from 2 August 2026 to 2 December 2027, and for high-risk AI embedded in products already covered by EU product-safety law to 2 August 2028.

What it did not do is change the substance. The obligations are intact; the enforcement date moved. That distinction is the one that gets lost in the retelling, and the practical consequence is specific: an organisation that reads the deferral as permission to stop work arrives at December 2027 needing to produce an assessment about a system it has by then been running for eighteen months, with no contemporaneous record of what it was doing in the interim. The assessment you cannot write is the retrospective one.

So treat the extra time as what it actually is — room to build the measurement side, which is the expensive part and the part you cannot backfill. The document is a week of work. The telemetry that makes the document true is a quarter, and it only produces history going forward.

If you deploy into the EU in a FRIA-triggering role, put one line in your agent registry entry now: the date you first put the system into use. Article 27 ties the notification to first use, and a deployment date reconstructed from a cloud bill two years later is the kind of detail that turns a routine filing into a finding. The registry entry is the right home for it — see agent inventory and registry.

STEP 3

Name the configuration, because "the system" is not a stable referent.

Here is the agent-specific problem in one observation. A conventional risk assessment describes software that changes when someone deploys it. An agent's behaviour is determined by at least six things, five of which change without a deploy and three of which can change without anyone at your company doing anything at all.

# What actually determines this agent's behaviour, by change cadence

model + version        vendor's schedule, not yours      # can change silently
system prompt          edited by whoever owns the prompt  # often weekly
tool list + scopes     grows as integrations land         # per sprint
retrieval corpus       whatever got indexed last night    # continuous
guardrail thresholds   tuned in response to incidents     # ad hoc
autonomy level         raised when the team gains trust   # rarely logged

# Of these, how many go through your change-control process?

   typically: one (the code), and it is the least behavioural

An assessment keyed to "the customer-service agent" is therefore keyed to nothing checkable. Key it to the six fields above, record their values at assessment time, and your re-assessment trigger becomes mechanical rather than a judgement call: a change to any of them is a change to the assessed system. Two of the six are worth special attention — a model version that the vendor can move under you is the problem catalogued in unpinned vendor defaults, and an autonomy level raised informally is the single change most likely to invalidate an assessment while leaving no record that it happened.

This is also where the assessment earns its keep internally rather than legally. The exercise of writing down what actually determines behaviour tends to surface a tool nobody remembered granting and a prompt nobody owns, and both of those are findings worth having regardless of which regulator is asking.

STEP 4

Make every conclusion falsifiable by a monitor.

This is the load-bearing recommendation on the page. A risk conclusion written as prose — "the risk of discriminatory outcomes is mitigated by human review of all adverse decisions" — cannot be checked, cannot expire, and cannot be invalidated by anything short of an incident. Write each conclusion as four fields instead, and the document becomes a set of standing claims your telemetry either supports or contradicts.

# One row per risk conclusion. If a row has no measurement, it is an opinion.

RISK          adverse decision issued without human review
CLAIM         every adverse outcome is reviewed before it is sent
MEASUREMENT   % of adverse-outcome events with a recorded reviewer id
THRESHOLD     < 100% over any rolling 24h window
OWNER         head of operations (named person, not a team)
ON BREACH     assessment marked stale; deployment to advisory-only
RE-ASSESS     any change to the six configuration fields

RISK          affected person cannot contest the outcome
CLAIM         an appeal route exists and is reachable in two steps
MEASUREMENT   appeals opened / adverse outcomes, by channel
THRESHOLD     ratio below the floor agreed with the business
OWNER         complaints lead
ON BREACH     the route is broken or unknown; investigate within 5 days

The test of whether you have done this well is blunt: can a monitoring system invalidate this assessment without a human reading it? If no, the document is a snapshot and will be stale before it is filed. If yes, it is a contract between the risk function and the telemetry, and staleness becomes an alert rather than a discovery during an audit.

Two practical notes. The measurements should come out of the pipeline you already run for quality-regression detection, not a parallel compliance pipeline — a second set of numbers maintained by a different team diverges within a quarter and then nobody trusts either. And a breach of a threshold is not automatically an incident: the documented response should usually be "the assessment is stale, reduce autonomy, re-assess", which is a cheaper action than the reporting path in serious incident reporting and keeps that path for what it is for.

STEP 5

Article 27's elements, and where each one breaks on an agent.

Article 27 asks a deployer to describe the processes in which the system will be used, the period and frequency of intended use, the categories of natural persons and groups likely to be affected, the specific risks of harm to those categories, the human-oversight measures, and the measures to take if the risks materialise — including governance arrangements and complaint mechanisms. Each of those lands awkwardly on an agent, in a way that is worth anticipating.

  • Period and frequency of intended use. Written for a system someone switches on to make a decision. An always-on agent has no sessions, so state a rate and a ceiling instead — decisions per day, maximum concurrent runs, and the cap you actually enforce. A frequency you cannot enforce is not a description, it is a hope.
  • Categories of persons likely to be affected. The trap is that an agent affects people who never interact with it. Everyone named in a document it reads, everyone in a record it writes, everyone downstream of an action it takes. Enumerate from the agent's data and action surface, not from its user list.
  • Human oversight measures. This has to name a person who can actually stop the system, with the mechanism and the time it takes. If the honest answer is that stopping it requires a deploy, say so — that is a finding, and it belongs with kill switches rather than being papered over with a policy sentence. The design side of the same question is human in the loop.
  • Complaint mechanisms. For an agent, the thing that breaks is not the existence of a route but its reachability: an affected person has to be able to find out that an agent was involved at all. That is the contestability and appeals problem, and it is the element most often written as aspiration.
  • Governance arrangements. Name the role, not the committee. The useful version of this field is a single accountable person per deployment, as argued in accountability and roles.

Where the DPIA already answers one of these for the same processing, say so explicitly and cross-reference it rather than restating it in different words. Two documents that describe the same safeguard slightly differently is a worse position than one document and a pointer, and the inconsistency is exactly what a reviewer notices.

STEP 6

Where it lives, and what to do this quarter.

The operating model matters more than the template, and it comes down to four decisions.

  • Store it next to the registry entry, not in a document store. The assessment's configuration fields and the registry's configuration fields are the same fields. Keeping them in two places guarantees they disagree, and the disagreement will be discovered by someone external.
  • Version it with the configuration, not on a calendar. Annual review is the wrong cadence for something whose subject changes weekly. The trigger is a change to the six fields; the annual pass is a backstop for the case where nothing tripped.
  • Make "stale" a state the deployment can be in. Not a label on a document — a state with a defined consequence, usually reduced autonomy until re-assessment. This is the only mechanism that gives the assessment teeth between audits, and it is the same instrument as the degradation states in model risk management for agents.
  • Keep one person's name on it. Assessments owned by a function get reviewed by nobody.

Scope note: the assessment is not a substitute for the evidence it cites. Audit trails, retention and the lineage of what the agent read remain their own problems — see audit trails and provenance and data governance for agents. The assessment's job is to state what you believe and how you would know you were wrong.

Do this today: take your most consequential deployed agent and write down the six configuration fields from Step 3 with their current values. Then ask, for each one, who changed it last and whether that change is recorded anywhere. The fields you cannot answer are the real finding — not because the agent is non-compliant, but because any assessment you write about it is unfalsifiable, and an unfalsifiable assessment is worse than none: it creates a documented belief that nothing will ever contradict. Fixing the record costs days; discovering the gap during a filing costs a deployment.