Alphanumerics Over Voice

9 min read

V18
Playbook · Voice & Realtime Agents

Alphanumerics over voice: stop trying to hear the string.

Your transcription is excellent and your order lookups fail, because a booking reference is the one thing in the call with no language model behind it — every character is an independent coin flip, and exact match is the product of ten of them. That arithmetic is why better recognition only ever buys you percentage points: going from 97% to 99% per character, a large win nobody will hand you twice, still leaves one call in ten unable to find the order. The fixes that change the shape all do the same thing, which is to stop putting an arbitrary string through the audio channel at all.

STEP 1

The metric on your dashboard cannot see this failure.

Word error rate is computed over words, and words have a language model behind them: context, collocation, morphology, a prior that repairs a mangled phoneme before you ever see it. RK4T-92B has none of that. Each character is drawn from a 36-symbol set with a near-uniform prior, the acoustic evidence is a few hundred milliseconds of band-limited telephony audio, and nothing downstream can correct a mistake because every string is equally plausible.

Model the consequence directly, because it is the whole argument on this page. For a ten-character identifier at per-character accuracy p, exact match is p10:

# Exact-match rate for a 10-character identifier

  per-character 95%   ->   0.95^10  =  59.9%
  per-character 97%   ->   0.97^10  =  73.7%
  per-character 99%   ->   0.99^10  =  90.4%

# A 6-character reference is much kinder:

  per-character 97%   ->   0.97^6   =  83.3%

# Doubling accuracy per character does not halve the failure rate.
# Removing characters does.

Published vendor figures land where the arithmetic predicts. One speech vendor reports its alphanumeric-tuned mode at 98.0% sequence accuracy on pure digit strings and 85.4% on mixed alphanumerics — the mixed case is the hard one, and the mixed case is what every order ID, SKU and booking reference actually is. The same write-up cites a comparison where 96.6% word-level accuracy sat alongside 77% exact match on the identifier itself: an excellent transcript and a failed lookup, in the same call.

Change the metric first, because you cannot manage this until you can see it. Score field-level exact match at first attempt, per field type, and treat word error rate as a diagnostic for the conversational half of the call only. The wider case for outcome-shaped scoring is in evaluating voice agents; this page is what to do once the number is honest.

STEP 2

Do the three cheap recognition fixes, and know their ceiling.

These are real improvements, they take an afternoon, and none of them changes the exponent:

  • Turn on the provider's alphanumeric or spelling mode, per turn. Most streaming STT vendors ship a mode that suppresses the word-level language model and decodes character by character. Enable it only for the turn where you expect an identifier — leaving it on for the whole call degrades the conversational half badly.
  • Bias toward the strings that can exist. Keyword boosting, phrase hints and custom vocabulary lists are supported almost everywhere and are dramatically effective when the candidate set is small. If you know the caller's three open orders, boost those three strings.
  • Constrain the decode to your actual format. If your references are always four letters, a hyphen and three digits, do not accept a decode that is not. Rejecting the impossible is free accuracy, and it converts a silent wrong answer into a re-ask you can handle.

Then know what you are still fighting. Letter names are acoustically confusable in a way that no amount of tuning removes: the E-set — B, C, D, E, G, P, T, V and Z — differs only in a short burst before a shared vowel, and over an 8 kHz telephony channel the cues that separate them are largely gone before the model sees anything. M and N, F and S, and A against the digit 8 fail the same way. The channel constraint is structural, and the speech stack entry explains why 8 kHz changes everything upstream of your choices.

STEP 3

The move that actually works: resolve against a set you already have.

This is the recommendation the rest of the page exists to support. Almost every identifier a caller reads aloud is one you could have looked up another way, and you already hold a small candidate set — the orders on this phone number, the bookings under this account, the three open tickets for this customer. Once the set is small, you are no longer transcribing an arbitrary string; you are ranking a handful of known strings against noisy evidence, and a two-character match is often decisive.

What that looks like in the call:

  • Identify the caller first, then the object. ANI, the account they authenticated against, an email on file. The identifier becomes a disambiguator between two or three candidates rather than a primary key. See caller authentication for doing the first half without turning the agent into an oracle.
  • Match fuzzily against the candidates, not exactly against the alphabet. Score each candidate with an edit distance that is weighted by the known confusions — treat B/D/P as nearly free substitutions, treat a vowel change as expensive — and accept when one candidate is clearly ahead.
  • Ask a describing question instead of a spelling question. "Is that the order from Tuesday, or the one going to the Manchester address?" resolves two candidates in one turn with no character-level recognition at all, and it is the turn callers find easiest in the entire interaction.
  • Only fall back to capture when the set is genuinely unbounded — a first-time caller, a reference from another company, a document number you have never seen.

The reframing worth internalising: a 3610 search space that you were trying to hear is usually a 1-of-3 choice that you already had the data to make. Teams reach for a better STT vendor when the actual fix is a database query they were running one turn too late.

STEP 4

When you must capture a free string, use a protocol, not a prompt.

Sometimes the set really is unbounded. Then the capture is an interaction design problem, and the design that works is not "ask, listen, confirm at the end":

  • Chunk into groups of three or four, and confirm each group. "Let's take it in pieces — the first four?" An error inside a confirmed chunk costs one re-ask of four characters. An error inside a ten-character read confirmed at the end costs the whole string, and callers re-read it differently the second time, which produces a new error in a new place.
  • Read back with disambiguation, one direction only. "B as in bravo, four, seven, D as in delta." Offer the phonetic alphabet in your own readback; never require the caller to use it. Asking a caller to spell phonetically is the single most reliable way to make an ordinary person hang up.
  • Never ask for the whole thing again. "Sorry, can you repeat that?" is the failure mode that turns one bad turn into an abandoned call. Re-ask the specific chunk, and say which one: "I have B-4-7; I missed the next two."
  • Treat DTMF as an equal path, not a punishment. "You can also type it on your keypad" offered before the first failure converts far better than the same sentence offered after the second. Digits over DTMF are near-perfect; letters are not, which is another argument for the format decision in Step 6.
  • Send a link when there is a data channel. If the call has an accompanying app or SMS session, a tap beats any amount of speech. This is not a voice failure; it is using the channel that fits the payload.

Payment card numbers are not an alphanumeric capture problem and must not be solved with the protocol above. Card data must never reach the agent leg at all — the DTMF suppression and pause-and-resume mechanics, and the reason the buffer usually starts too early, are in recording consent and redaction.

STEP 5

Set the confirmation bar by consequence, not by confidence.

The instinct is to confirm everything above some recognition-confidence threshold. That is the wrong axis: the confidence score tells you about the audio, and what you need to know is what happens if you are wrong. Sort the fields by the cost of a false accept:

  • Irreversible or financial — confirm explicitly, every time. The account a refund lands in, an address a replacement ships to, an amount. One extra turn is cheap against a wrong-account credit.
  • Reversible and visible — resolve probabilistically and state what you did. "I've pulled up the order ending 7-4-D, shipping to Manchester." The caller corrects it immediately if it is wrong, which is a cheaper confirmation than asking.
  • Read-only and low-stakes — do not confirm at all. Reading out a delivery date for the wrong order is a mistake the caller catches for free. Confirming it costs every caller a turn to protect against a harmless error.
  • Never confirm by repeating a digit string the caller must verify in their head. Callers say yes to things they did not fully parse, especially under time pressure. Confirm by naming a property they recognise — the date, the destination, the item — rather than by reciting the key back to them.

Numbers spoken back carry a separate normalisation problem per language — "double seven", "seventy-seven", grouping conventions — which is handled in multilingual voice agents and is a real source of confirmations that read correctly and mean something else.

STEP 6

Measure capture as its own funnel — then go change the identifier.

Four numbers, per field type, tell you whether any of the above worked:

  • First-attempt field exact match. The headline. Segment it by channel — mobile, landline, VoIP, speakerphone — because the spread is wide and the fix differs.
  • Turns to capture. The caller-experience number. Three turns to read one reference is the point at which people start asking for a human, and it often precedes abandonment by one turn.
  • DTMF fallback rate, split by whether it was offered or requested. Offered-and-taken is a success. Requested-after-failure is a design defect with a timestamp on it.
  • Abandonment after the second re-ask. The cliff. If it is steep, your re-ask strategy is asking for the whole string again; go back to Step 4.

Then take the finding upstream, because the most effective intervention is not in the voice stack at all. Identifier formats are chosen by whoever built the ordering system, usually with no idea that a human would ever read one aloud, and they are changeable:

  • Drop the confusable letters. A Crockford-style alphabet that excludes I, L, O and U removes the worst visual collisions; excluding the E-set consonants removes the worst acoustic ones. You lose a little entropy per character and buy it back with one more character.
  • Add a check character. A single checksum digit turns most single-character misrecognitions into a detected error rather than a wrong lookup — and a detected error can be re-asked precisely instead of silently returning the wrong order.
  • Shorten it, or split it. Six characters at 97% beats ten at 99%, and a short public-facing reference mapped to a long internal key costs one table.

If you do one thing: pull fifty real calls where the agent asked for an identifier and score first-attempt field exact match by hand. It will be well below what your word error rate implied, and the gap is the entire business case. Then, before touching the speech stack, check whether the caller's phone number already narrows that identifier to fewer than five candidates — for most support and retail flows it does, and the whole problem turns into a disambiguation question you can ask in one friendly turn. Related: voice agent failure modes for what the surrounding call does when this goes wrong, replacing an IVR for why the keypad path is worth keeping, and escalation and warm transfer for handing over without making the caller read it a third time.