Contestability & Appeals

8 min read

C19
Operation · Governance & Compliance

Contestability: an appeal you cannot reconstruct is a coin toss with paperwork.

The appeal arrives six weeks after the decision, and by then the model is two versions on, the retrieval index has been rebuilt, the policy document has been edited and the prompt has been tuned twice. Whatever you re-run is a different system answering a different question, so the reviewer is not checking a decision — they are making a fresh one and calling it a review. Contestability is not a form and a mailbox; it is the ability to put the original inputs back in front of a person who has the authority to overturn. That capability is bought at decision time, in your retention policy, or it is not available at all.

STEP 1

Three different rights, routinely collapsed into one.

Teams build "an appeals process" and discover late that the obligations are separable, arrive from different instruments, and land on different parties.

  • The right to human intervention — GDPR Article 22(3), where a decision based solely on automated processing produces legal or similarly significant effects. The data subject may obtain human intervention, express their point of view, and contest the decision. Note the trigger: solely automated. A human who rubber-stamps does not remove you from Article 22; a human with real authority and information does.
  • The right to an explanation of an individual decision — EU AI Act Article 86, owed by the deployer (not the provider) to a person affected by a decision taken on the basis of a listed Annex III high-risk system, excluding point 2, where the decision has legal effects or similarly significantly affects them adversely. It is reactive and on demand: clear and meaningful explanation of the role of the AI system in the procedure and the main elements of the decision taken.
  • The right to contest and have the decision reconsidered — the substantive one, and the only one that changes an outcome. The first two are inputs to it.

These are cumulative, not alternatives, and the sector rules you already live under — lending, insurance, employment, benefits, healthcare — usually add their own deadlines and notice contents on top. Map them once, per decision type, and record which ones apply where; the mapping belongs beside your agent registry rather than in a slide deck.

The AI Act obligation falling on the deployer is the detail that surprises people. If you bought the model or the agent, the explanation is still yours to give, from evidence the vendor may not expose. Make "can we explain an individual decision from data we hold" a procurement question, not a discovery you make after the first complaint.

STEP 2

Your retention policy is your appeals policy.

An agentic decision is not a row in a table; it is the output of a pipeline whose every component drifts. To place the original decision in front of a reviewer you must pin all of it, at the moment it happens, because none of it is recoverable afterwards.

# Pin at decision time. None of these can be reconstructed later.

decision_id        # the join key for everything below
subject_ref        # who it was about, and under which case
model_id + version # the exact deployed build, not "gpt-5-ish"
prompt_version     # hash of system prompt + tool definitions
policy_version     # the rules doc as it read that day
retrieved_docs     # ids AND content hashes — the index gets rebuilt
tool_calls         # arguments and results, not just names
inputs_as_supplied # the applicant's data before any normalisation
outcome + reasons  # the decision, and the grounds given to the subject
human_step         # who reviewed, what they saw, what they could change

# Retention: the longest of the appeal window, the limitation
# period, and the sector rule. Usually years, not the 30 days
# your trace backend defaults to.

Two of these lines are where systems fail. Retrieved content hashes, because a reindex or a document edit silently changes what "the same query" returns, and a reviewer comparing against today's corpus will not see what the model saw — the migration hazard in reindexing and embedding migrations. And retention length, because trace tooling is priced and configured for debugging, with sampling and a short window, while an appeal window is measured in months and a limitation period in years. Sampled traces are fine for observability and fatal for contestability: the appealed decision is, by definition, not the one you kept at random.

So carve out a separate, unsampled, long-retention path for decisions with legal effect, and note that it now interacts with retention and legal hold in both directions — you must keep it long enough to answer, and you must be able to place a hold on it when a case turns contentious.

STEP 3

Re-running the model is not a review.

The tempting implementation is a button that re-executes the pipeline on the original inputs and reports whether the outcome changed. It is cheap, it produces a document, and it is close to worthless — because the second run is not independent of the first. It shares the model, the prompt, the retrieval corpus and the framing, so it reproduces the original error whenever the error came from any of those, which is most of the time.

It gets worse when a person is in the loop badly. Show the reviewer the model's decision and its stated confidence first, and you have built a machine for manufacturing agreement: sycophancy pulls a model reviewer toward whatever is already on the table, and automation bias pulls a human reviewer the same way. Either way the second opinion correlates with the first, and correlated opinions do not average out — the point argued in the generator–verifier gap.

  • Give the reviewer a different information basis, not a re-run. The applicant's new evidence, the source documents rather than the model's summary of them, the policy text itself. If the only new input is "the model was asked again", nothing was reviewed.
  • Withhold the model's confidence, and consider withholding its recommendation, until the reviewer has formed a view. Anchoring is the whole mechanism of automation bias, and ordering is free to change.
  • The reviewer must be able to overturn without asking permission. Article 22(3) is not satisfied by someone who can only escalate. Authority, competence in the domain, and time are the three things that make review meaningful, and the third is the one quietly removed by a queue target.
  • Watch the overturn rate as a control, not an outcome. A rate near zero means the review is ceremonial; a rate near half means the primary decision is broken. Neither is a number to celebrate, and the first is the one that gets reported as success.
STEP 4

Design the notice, because it determines whether anyone can appeal at all.

Most appeal channels fail before the queue, at the decision notice. A person who cannot tell what the decision turned on cannot contest it, and a right that requires the subject to guess the grounds is a right in name.

  • Say that an automated system was involved, and what it did. The AI Act asks for the role of the system in the procedure — screened, scored, recommended, decided — in language a non-specialist reads once.
  • Give the main elements the decision turned on, not a feature-importance chart. "Your reported income for March–May could not be matched to the payslips supplied" is contestable. "Risk score 0.83" is not.
  • Name the route, the deadline and what evidence helps. Telling someone what would change the answer is the single highest-leverage sentence in the notice, and it converts a large share of appeals into a document upload instead of a case.
  • Do not make the appeal harder than the application. If the decision was made in nine seconds through an app and the appeal requires a posted form, you have built a filter, and a regulator will read it as one.
  • Log the notice you actually sent, versioned. What the subject was told is itself evidence, and "we would have said" is not a defence.

Where the decision is delivered by an agent conversationally, the disclosure and the grounds have to survive the conversation — the transcript is the notice, and it needs the same versioning as a letter. The interaction design side of this is transparency and explainability.

STEP 5

Overturning is a side-effect problem, not a status change.

The decision did not sit still while the appeal was pending. An agentic pipeline acted on it: it sent letters, cancelled a benefit, opened a collections case, wrote to a third party, updated a score another model consumed. Setting a flag from declined to approved reverses one row and none of that.

  • Enumerate downstream effects per decision type in advance, and store them with the decision record. You cannot compensate for an action you cannot list, and the list is short enough to write down once — this is the ledger that repairing agent side effects is built around.
  • Notify the third parties you told. A credit reference, a partner, another agency. Reversal without retraction leaves the original decision in circulation, which in several regimes is its own breach.
  • Make the subject whole on timing, not just outcome. Backdate the entitlement, refund the fee, remove the interest accrued during the dispute. An appeal that takes eleven weeks and restores nothing but status has resolved the record, not the harm.
  • Purge the reversed decision from memory and training paths. An overturned outcome that stays in a memory store, a cached summary or a fine-tuning corpus will be repeated, and the second time it will look consistent.
STEP 6

The appeal queue is the highest-quality evaluation data you will ever get.

Everywhere else you pay for labels. Here, motivated humans identify your system's errors, supply the missing evidence, and a qualified reviewer adjudicates them — for free, in a structured queue, continuously. Teams that treat appeals purely as a compliance cost throw that away and then commission a labelling project to reconstruct it.

  • Code every overturn by reason, using a fixed taxonomy: bad input data, missing evidence the subject held, policy misread, retrieval miss, model error, edge case the policy never addressed. That distribution tells you what to fix, and it is the failure taxonomy you would otherwise have invented in a workshop.
  • Promote overturned cases into the regression suite, with the correct outcome as the label. This is the cleanest eval set in the building, and it grows on its own.
  • Alert on appeal-rate and overturn-rate by segment. A rise in one cohort is a fairness signal and a drift signal at once, and it usually moves weeks before an aggregate quality metric does — a production feedback signal with a name attached.
  • Report time-to-resolution as a service level. Appeals decay: evidence goes stale, harm compounds, and a right exercised eleven weeks late has partly expired. This is the number that tells you whether the channel is real.

Start with the pin list in STEP 2 and one number. Write the ten fields at decision time for every decision with legal effect, on an unsampled path retained for the longest of your appeal window, limitation period and sector rule — that single change is what makes every later obligation answerable, and it is worthless if added afterwards. Then instrument the overturn rate and read it as a control: near zero means your human review is ceremonial and Article 22 is not satisfied, however many reviewers you employ. Everything else on this page — the notice, the side-effect ledger, the eval set — is buildable later. The record is not.

Related: audit trails for the record this hangs off, the EU AI Act for agents for where Article 86 sits among the deployer duties, accountability and roles for who signs the overturn, and public benefits casework agents for a domain where every line of this is already statutory.