Accessible Voice Agents

9 min read

V17
Playbook · Voice & Realtime Agents

Accessible voice agents: the average hides the callers who cannot go around you.

Your voice agent does not have a failure rate; it has a different failure rate for every kind of voice, and the single number on your dashboard is an average over a caller population you never stratified. The cohorts that fail hardest — disordered or slow speech, strong accents, older callers, anyone on a relay service — are the ones with the fewest alternatives, so their failures arrive as hang-ups rather than complaints and never enter your metrics. Worse, the standard turn-taking design is not merely blind to them: silence-based endpointing is itself the mechanism that produces the failure. Stratify the metric, make the endpointer adaptive, and keep two paths open that do not require speech.

STEP 1

Stratify the metric before you argue about anything else.

Containment, resolution rate, CSAT and word error rate are all averages, and an average is exactly the wrong statistic for a failure that concentrates. A 6% task-failure rate is consistent with 3% across most of your callers and 40% in a cohort that is 5% of volume — and the second reading is the one that describes a service people cannot use.

  • You cannot stratify by disability, and you should not try. Do not ask, do not infer, do not store a guess. What you can do is stratify by acoustic and interactional properties you already measure on every call, which correlate with the cohorts that struggle without ever labelling a person.
  • Five proxies that need no new data. Speech rate in words per minute; the distribution of intra-utterance pause lengths; number of ASR re-prompts in the call; whether DTMF fallback was used; and whether the caller hung up within ten seconds of a re-prompt. Bucket by each and read your existing success metric per bucket.
  • The abandonment signal is the one that matters. A caller who fails and escalates shows up in your transfer rate. A caller who fails and hangs up shows up nowhere, and the second behaviour is far more likely from someone who has been failed by phone systems before. Count hang-ups that follow a re-prompt as a distinct outcome, not as a generic abandon.
  • Report the ratio, not the cohort. The number to put in front of a product owner is worst-bucket success divided by median-bucket success. A single figure, comparable across releases, and it moves when you fix something.

Expect the first measurement to be uncomfortable and expect the slowest-speech bucket to be the worst one by a distance. That bucket contains people with speech disorders, people recovering from a stroke, people in their eighties, people using an augmentative communication device, and people composing a sentence in their second language. They have almost nothing in common except that your endpointer treats all of them the same way.

STEP 2

Recognition is the layer that breaks, and the disparity is documented.

The model is rarely the problem. The speech stack is, and its error rates are not uniformly distributed — a point with a long measurement record rather than a speculative one.

  • The 2020 Koenecke study of five major commercial ASR systems found an average word error rate of 35% for Black speakers against 19% for white speakers on matched conversational material. A near-doubling, in production systems, on the same task.
  • Work published in 2026 on three major commercial engines found substantially degraded caption accuracy for accented speech relative to a "standard" broadcast accent, on both word error rate and a semantic-similarity measure — the second metric mattering because it captures errors that change meaning rather than merely words.
  • For disordered speech, studies consistently find a large gap between read and conversational speech that widens with severity, and that gap persists in speaker-personalised models. Personalisation helps; it does not close it, and it is unavailable to you on an inbound call from a stranger.

The reason this is an agent problem and not a transcription problem is what happens next. Your agent does not read the transcript — it extracts tool arguments from it. A recognition disparity becomes an action disparity: the wrong account looked up, the wrong amount confirmed, the wrong branch taken. Multilingual and code-switching agents covers the neighbouring case where the damage lands on names and numbers; here it lands on whether the caller is understood at all.

STEP 3

Silence-based endpointing is the discrimination mechanism.

This is the part that is genuinely a design defect rather than a model limitation. A typical endpointer declares the turn over after 500–800ms of silence — a threshold tuned on fluent speakers, where it produces the snappy turn-taking your latency budget is built around.

Now run a caller whose pauses are longer than that through the same loop. They pause mid-sentence to breathe, to find a word, to press keys on a communication device, or because a stammer blocked. The endpointer fires. The agent starts speaking over them. They resume; barge-in cuts the agent off; the agent now has a fragment, cannot parse it, and re-prompts — often with a shorter prompt, which makes the next window feel even more hurried. This is the infinite apology loop in voice failure modes, arriving deterministically rather than occasionally, for an identifiable group of people.

  • Make the threshold adaptive and sticky. On the first re-prompt, raise the endpointing threshold substantially — 1,500ms is a reasonable starting point — and keep it raised for the remainder of the call. Never let it ratchet back down. The cost is a slightly slower conversation for a caller who was already struggling, which is the correct trade.
  • Prefer semantic end-of-turn over silence. A model that decides the turn ended because the utterance is syntactically complete tolerates pauses that a timer cannot. See turn-taking and barge-in; this is the strongest single accessibility argument for that architecture, and it is rarely the one it is sold on.
  • Lengthen the prompt on retry, do not shorten it. The common instinct — terse re-prompts to save time — removes exactly the context a struggling caller needs and signals impatience. Re-prompt with a concrete example of the expected answer.
  • Suppress barge-in during confirmations only. Barge-in is mandatory almost everywhere. The exception is a confirmation of an irreversible action, where a cough should not commit a payment.
  • Watch for the cascade with full-duplex stacks. Removing the turn removes the timer, which helps — and introduces a new failure where the agent backchannels over a long pause and the caller stops to listen. Test it against slow speech specifically.
STEP 4

Two paths that never require speech, and one channel you have probably not tested.

Every accessible voice deployment keeps a non-speech route open at all times, not as a menu item the caller had to hear and remember.

  • DTMF live at every prompt, not only at the top. Keypad entry is the most reliable input channel available to a caller whose speech is not being recognised, and it is usually implemented as a one-time menu option and then switched off. Accept digits at any point in the conversation, and say so once in the opening.
  • A human path reachable without a magic phrase. "Say representative" fails precisely for the callers who need it. Zero should always work, three consecutive re-prompts should escalate automatically, and the escalation should carry the context packet described in escalation and warm transfer — a caller who has already repeated themselves four times must not start over.
  • Relay services break every timing assumption you have. On a TTY, IP or captioned-telephone relay call, a human or automated intermediary sits in the middle: silences run to tens of seconds, the speaker is not the caller, phrasing arrives in the third person ("she says her account number is…"), and audio quality is whatever the relay leg gives you. An agent that endpoints on silence, authenticates on voice, or assumes first-person phrasing fails these calls categorically. Test at least one relay path before launch; most teams have never placed the call.
  • Do not make speech the only authentication factor. Voiceprints exclude people whose voice varies with their condition, and asking a caller to speak a long identifier is precisely the task recognition handles worst. Caller authentication argues for moving the proof out of the audio channel on security grounds; accessibility is the second, independent reason to do it.
STEP 5

The output side, which is cheaper to fix and usually ignored.

Half of accessibility on a voice channel is whether the caller can take in what the agent said, and none of it involves the recogniser.

  • Never deliver critical information exactly once. A reference number, an amount or an appointment time should be repeated, chunked ("four-two-one … eight-nine-three"), and confirmable on request. "Say that again" must work at any point and must not restart the turn.
  • Cap spoken options at three. A spoken list is held in working memory, and the fourth item pushes the first one out. If you have seven branches, ask a disambiguating question instead of reading seven options — this is also the fix for the menu tree you inherited in replacing an IVR.
  • Make TTS rate adjustable and start slower than your demo. Voice designers tune speech rate on themselves. A rate that sounds crisp to the team is fast for an older caller and fast for anyone listening in a second language. Offer "slow down" as a recognised command.
  • Do not let audio be the only copy. Where you have a channel — SMS, email, an account page — send the reference number there too. It costs nothing and it converts a comprehension failure into a non-event.
  • Mind the silence. Long gaps during tool calls read as a dropped call and cause hang-ups. Preambles are covered in voice tooling and state, but the accessibility-specific point is that the caller who hangs up during dead air is usually the one who has already had a hard call.
STEP 6

Make it a release gate, because a report changes nothing.

Everything above degrades back to the default the first time someone tunes the endpointer for latency. The only durable version of this work is a number in CI that blocks a deploy.

  • Build the cohort set from real recorded audio. Synthesised "accented speech" is not a test of anything — TTS produces fluent, well-formed audio with an accent painted on, and fluency is the variable that matters. Source consented recordings across speech rates, disfluency, accent, age and line quality, and treat that set as the asset it is. Evaluating voice agents is where the harness comes from; this adds the stratification.
  • Gate on the ratio. Fail the build when worst-cohort task success drops below a fixed fraction of median-cohort success. A ratio gate survives a general quality improvement, which an absolute threshold does not.
  • Score timing separately. Endpointing errors — cut off mid-utterance, or waited too long — are their own metric, per cohort. They are the leading indicator; task failure is the lagging one.
  • Keep the escalation path in the gate. Assert that three consecutive re-prompts reach a human, in the test suite. This is the control that bounds the damage of everything you have not thought of, and it is the first thing a refactor breaks.
  • Log the complaint route. A caller who could not use the agent needs a way to say so that is not the agent. See contestability and appeals.

Do these three this month and you will have most of the available benefit: bucket last month's calls by speech rate and read your existing success metric per bucket; raise the endpointing threshold to 1,500ms permanently after the first re-prompt; and place one test call through a relay service yourself. The first tells you the size of the problem, the second fixes a meaningful share of it for the cost of one config change, and the third will surface a failure that no amount of transcript review would have found. Related: accessible agent interfaces for the screen equivalent, turn-taking and barge-in for the mechanism, and evaluating voice agents for the harness this plugs into.