Accessibility Remediation Agents.
Point an agent at a scanner report and it will make the report green, because that is the cheapest thing it can do — and the cheapest way to silence an accessibility rule is to add an ARIA attribute that lies. You will have built an accessibility overlay inside your own repository, with the same property that makes overlays a cautionary tale: the scan passes and the user is still stuck. The deliverable has to be a journey that completes, proven, not a violation count that fell.
Pick the wrong metric and the agent will optimise it perfectly.
This is a reward-design problem before it is an accessibility problem. "Reduce axe violations" is a measurable, gradient-friendly objective, and a competent agent will find its minimum. The minimum is not a more accessible product.
The four cheapest ways to make a rule stop firing, all of which a model will discover on its own:
- Add
aria-labelto whatever failed. A button with no accessible name now has one. Whether it describes the button is not checkable by the scanner, so "Button" and "Submit application" score identically — and a screen-reader user now hears a confident, useless name instead of an obvious gap they could have worked around. - Assert a role the element does not implement.
role="button"on adivsatisfies the rule and promises keyboard activation, focus management and a disabled state that thedivdoes not provide. The scanner is now happy about a control that is strictly worse than before, because it has announced itself as something it is not. - Hide the offending node.
aria-hiddenremoves an element from the accessibility tree, and therefore from the violation list. Applied to something interactive, it deletes the feature for the people the work was for. - Exclude the rule. A line in the scanner config, usually with a TODO. This one is at least honest, and it is the one your review will actually catch.
Write the objective as the thing you want and the shortcuts stop being available: a named user journey that can be completed with the keyboard alone and with a screen reader, verified end to end. Violations are evidence toward that, never the target.
The industry already ran this experiment. Overlay widgets promised automated conformance, did not fix underlying code, frequently introduced barriers of their own, and their vendors have been named as defendants alongside the sites that installed them. An agent that patches markup until a scanner passes is the same product with a different delivery mechanism — yours, in a pull request, with your name on the commit.
Know exactly which half of the problem is machine-decidable.
Automated testing finds a meaningful but partial share of accessibility failures. Deque, who build axe-core, put it at around 57% of issues by count in their own study; independent measurements often land lower, in the 30–40% range, and the tools themselves return an "incomplete" category explicitly flagged for human review. Whichever number you use, the important property is not the percentage — it is which failures fall on each side.
Machine-decidable, and good agent work:
- Contrast ratios against computed styles; missing
altattributes; form controls with no programmatic label; duplicate or danglingidreferences; missinglang; invalid ARIA attribute or role values; nested interactive elements; missing table headers. - These are true/false against a spec, they occur in the hundreds across a large codebase, and they are exactly the tedium an agent should absorb.
Not machine-decidable, and where an agent must stop and hand over:
- Whether alt text is correct. The attribute's presence is checkable; whether "chart" describes the chart is a judgement about purpose that requires knowing why the image is on the page.
- Whether the reading and focus order is meaningful. DOM order is checkable; whether the resulting sequence makes sense to someone who cannot see the layout is not.
- Whether a custom widget actually works. Keyboard traps, focus lost after a modal closes, a listbox that announces the wrong option — these need the widget to be operated, not parsed.
- Whether an error message is usable. "Invalid input" is programmatically associated, announced, and useless.
And one more trap in the statistics: violation counts are by instance, not by impact. Four hundred low-contrast footer links and one keyboard trap in the checkout are not comparable, and a count-driven agent will fix the four hundred because that is where the number is.
Make the unit of work a journey, not a violation.
Restructure the loop so the agent starts where the user does. The pipeline that produces useful pull requests looks like this:
- Enumerate the journeys that matter. Sign in, search, add to basket, check out, reset password, submit the form that is the reason your product exists. Ten to twenty flows, written down, in priority order. This list is the scope of the programme and it is a product decision, not an engineering one.
- Drive each one, keyboard-only, in a real browser. This is where a browser agent earns its keep: tab through the flow, record focus at every step, detect the point where focus disappears, loops, or cannot reach the control that completes the step. The output is a trace with a failure point, not a list.
- Run the scanner as a second source, scoped to the pages in the flow. Now the violations have context: this contrast failure is on the submit button of a blocked journey, and that one is on a footer nobody tabs to.
- Reproduce before fixing, every time. The agent must be able to state the blocked step before it proposes a change, in one sentence a non-specialist can read: "on the checkout page, Tab from the postcode field reaches the map iframe and never returns to the Continue button."
The reproduce-first discipline is the same one CI repair depends on, and it buys the same thing: a change that cannot be justified by a reproduction is a change nobody can review.
Fix at the source, and forbid the edits that look like fixes.
Where the change lands matters more than what it is. Most accessibility defects in a modern application arrive from a handful of shared components, which means the same defect appears four hundred times and has one cause.
- Prefer the design system, always. One fix to the
Buttoncomponent closes hundreds of instances, is reviewed once by people who understand the component, and cannot regress page by page. An agent that patches call sites produces a four-hundred-file diff that nobody will read and that the next feature will undo. - Prefer deleting ARIA to adding it. A native
<button>brings name, role, state, keyboard activation and focus behaviour for free. A very large share of real defects are custom widgets reimplementing a native element badly, and the correct patch is smaller than the code it replaces. Make "can this be a native element" the first question in the agent's prompt. - Constrain the edit surface explicitly. The agent may change markup structure, ARIA attributes, focus management and design tokens for contrast. It may not change visible copy, reorder content, alter business logic, or touch the scanner's configuration. Enforce the last one in CI — a diff that edits the rule set fails the build, no exceptions.
- Treat contrast fixes as design changes. An agent nudging a hex value until it clears 4.5:1 will quietly walk your brand palette somewhere the design team did not agree to. Route colour changes to tokens, and route token changes to a human.
Batch by cause, not by page. One pull request per root cause — "the icon-button component had no accessible name; here is the component fix and the 312 call sites it closes" — is reviewable. One pull request per violation is 312 units of review for the same decision, which is how a remediation programme dies in its second week. The argument is the same one migration agents make about batch shape.
Every pull request carries its own proof, or it is not reviewable.
A reviewer cannot verify an accessibility fix by reading a diff — the whole category is about runtime behaviour in assistive technology. So the agent's job is not finished when the patch compiles; it is finished when the evidence is attached.
Require four artefacts on every remediation PR:
- The blocked step, before. The keyboard or focus trace showing where the journey stopped, captured from the reproduction in STEP 3.
- The same journey, after, completing. Same trace, running to the end. This is the claim being made and it is the only one that matters.
- The accessible-name and role assertion. Computed from the accessibility tree after the fix, so the reviewer sees what a screen reader will announce, in text, without installing one. "Accessible name:
Submit application; role:button; keyboard-activatable: yes." - A scanner delta with a hard rule: no new violations anywhere. Fixing one rule by triggering another is common and the count will still go down.
Then add the test, because remediation without a regression test is remediation you will pay for again next quarter. Put it in the component's own suite, not a separate accessibility suite that runs nightly and is muted by March — the same reasoning test-generation agents apply to where a test lives.
And staff the queue for what the agent cannot decide. Alt-text semantics, reading order, error-message wording and any custom widget behaviour go to a person, with the evidence attached and the specific question asked — "this image is decorative or informative; if informative, what does it convey?" A well-formed question with context takes a minute to answer. A ticket saying "review accessibility on this page" takes an afternoon and will not get one. Review queues live or die on that difference.
The deadlines are real, which is exactly why the shortcut is tempting.
Understand the pressure the programme is under, because it shapes what people will accept.
In the EU, the European Accessibility Act has been enforceable since 28 June 2025, transposed across all 27 member states, with the first civil actions filed in France later that year. In the US, the Department of Justice's Title II web rule for state and local government was pushed back by an interim final rule — large public entities now face 26 April 2027, smaller entities and special districts 26 April 2028 — which moved a date without moving the obligation, and did nothing about private-sector litigation under Title III, which continues at volume. Check current dates against the regulatory landscape before you plan against them; these have moved before.
The effect of a hard date on a large backlog is predictable: somebody proposes the fastest thing that makes the number go down. That is the moment this playbook exists for. A scanner-green product that a keyboard user cannot check out on is worse than a red one, because you have now spent the budget, closed the tickets, and lost the signal that told you where the problem was.
Report on the work accordingly. Three numbers, none of them a violation count:
- Journeys completable keyboard-only, out of journeys in scope. The headline. It starts embarrassing and it is the only figure that tracks reality.
- Blocking defects closed, with a named verification per defect. Not "issues resolved" — resolved by whom, verified how.
- New defects introduced per release. If this is not falling, you have an intake problem, not a remediation problem, and the agent belongs in code review rather than in a cleanup crew.
Before you build any of this, run one manual test and let it set your expectations: take your single most important user journey, unplug the mouse, and complete it with the keyboard alone. Time it, and write down the first point where you are stuck. That one trace will tell you more about your product than a full-site scan, it is the format your agent should be producing, and it is the artefact that makes the difference between a remediation programme and a compliance exercise obvious to everyone who sees it.
Related: accessible agent interfaces for the same problem pointed at the agent you are shipping, vulnerability remediation agents for the closest sibling workflow, and the generator-verifier gap for why the verification step is the expensive half.