Replacing an IVR: the menu tree is not the spec, and containment is not the target.
The two artefacts every IVR replacement starts from are both traps. The call-flow diagram documents what touch-tone could express, not what callers want — the real requirement lives in the transcripts of everyone who pressed 0. And containment, the metric the whole industry quotes, scores the caller who gave up as a success, which is why vendor numbers jump from single digits to eighty percent without anyone's problem being solved. Build from the zero-out corpus, migrate one intent at a time in front of the IVR you already trust, and measure resolution without a callback.
The tree tells you about 1998; the zero-out transcripts tell you about your callers.
An IVR menu is a compression of caller intent into the bandwidth of twelve keys. Every node is a decision someone made about which four options fit in a prompt a caller would tolerate hearing, and everything that did not fit was routed to "for all other enquiries, press 0" or to an agent queue. Rebuilding that tree in natural language reproduces a set of constraints that no longer exist and inherits every compromise made under them.
- Mine the zero-out and "other" paths first. Those calls are your unserved demand, already labelled by the caller's own behaviour. Pull six to twelve months of recordings and transcripts from that branch and cluster them; the top ten clusters are the roadmap.
- Mine the misroutes. Calls that reached a queue and were transferred once tell you where the tree's categories do not match reality. A category that gets 30% onward transfer is a category the caller could not map their problem onto.
- Mine the repeat contacts. A second call within 72 hours about the same account is a resolution failure the IVR recorded as a success twice.
- Keep the tree for one thing only: the compliance obligations bolted to it. Recording disclosures, language offers, emergency routing, regulatory scripts. Those are requirements; the menu structure is not.
Do this before choosing a platform. The intent inventory determines whether you need a full conversational agent or a much cheaper natural-language front door that routes to the same six destinations — and for a large share of enterprise IVRs the honest answer is the latter, which changes the budget by an order of magnitude.
Containment counts abandonment as a win. Pick a metric that cannot.
Containment is the share of calls that never reach a human. It is the number in every vendor deck, legacy IVRs are usually quoted in the 5–10% range against 60–90% for voice agents, and the comparison is close to meaningless because the metric is satisfied by three very different events: the caller was helped, the caller was routed to self-service and gave up, or the caller hung up in frustration. Two of the three are failures and all three increment the same counter. Worse, "containment", "deflection" and "resolution" are used interchangeably in marketing material and mean different things.
- Primary metric: resolution without a callback in 72 hours. The call ended, the caller's problem was in fact handled, and they did not come back through any channel. This is the only number that cannot be gamed by making it harder to reach a human, and you can compute it from data you already have.
- Guardrail metric: abandonment inside the agent. Hang-ups during the automated portion, split by where in the conversation they happened. A rising containment rate with rising abandonment is a system getting worse and reporting better.
- Guardrail metric: repeat-contact rate and channel spillover. Calls "contained" that reappear as chats, emails or web tickets have not been resolved, they have been moved to a queue you were not looking at.
- Report handle time separately for the human tail. If the agent takes the easy calls, average handle time for the remaining human calls goes up mechanically. Comparing it to last year's figure will make your contact centre look like it regressed.
Fix the definitions in writing before the pilot, because the first executive summary sets the baseline everyone argues from for two years. The framing here is the same one evaluating voice agents applies at the turn level, lifted to the call.
Migrate by intent, in front of the IVR you are replacing.
The failure mode of an IVR replacement is a big-bang cutover on a Monday. The pattern that works puts the agent in front of the existing system and lets the old system remain the fallback, so every unhandled case lands somewhere already proven rather than somewhere new.
- Answer first, hand back on anything not yet migrated. The agent takes the call, opens with an open question, and for any intent outside its scope transfers into the legacy IVR at the correct node or straight to the queue. Your rollback is a config change, not a deployment.
- Rank intents by volume × handle time × automatability, and start with the boring high-volume one — balance enquiries, appointment moves, order status. The complex intent that excites the steering committee is the wrong first migration because it produces ambiguous evidence.
- Ramp per intent and per phone number, not per percentage of all traffic. A 10% ramp across everything gives you a thin sample of every failure mode at once; 100% of one intent on one number gives you a readable result in a week. The general discipline is in rollout and versioning.
- Never remove the path to a human, and honour it on the first ask. Beyond being decent, in several jurisdictions and sectors it is a requirement, and every hidden zero-out shows up in your abandonment number anyway.
- Freeze during peaks. No migrations during your seasonal spike, month-end, or a known incident — the periods when the fallback is under most pressure are exactly when you need it unchanged. Have a documented "revert to IVR" switch and rehearse it, per graceful degradation.
The integrations are the project; the conversation is the demo.
An IVR reads. It looks up a balance, a delivery date, a queue position — usually read-only, usually through a CTI adapter or a host connector built a decade ago. An agent that actually resolves things has to write: reschedule the appointment, apply the credit, update the address. That is the whole difference in effort, and it is invisible in every demo.
- Enumerate the writes before you scope the conversation. Each one needs an idempotency key, a confirmation read back to the caller by ear, and a compensating action if the call drops after the write and before the confirmation. The state discipline is voice tooling and state.
- Authenticate to the level the action needs, not the call. Reading a delivery date and changing a bank detail are not the same risk, and voice is the channel where the impersonation economics are worst — see caller authentication.
- Keep DTMF where DTMF is better. Account numbers, card digits, dates of birth, PINs. Keypad entry is more accurate than speech on 8 kHz telephony audio, it is what your PCI scope already assumes, and callers in public places will not say a card number aloud. A voice agent that cannot fall back to digits is a downgrade.
- Budget for the audio, not the transcript. Narrowband telephony degrades recognition on exactly the tokens that matter — names, postcodes, alphanumerics — and every accuracy figure you were shown was measured on wideband audio. The speech stack and latency budget pages carry the numbers.
- Carry the disposition through. The IVR wrote a routing code that downstream reporting depends on. The agent must write an equivalent, or six dashboards break silently — see after-call work and CRM writeback.
What the IVR gave you free, and what you now owe.
A menu tree is a deterministic program. Everything it did was inspectable, repeatable and printable — properties nobody valued until they were gone, and several of them are contractual.
- A printable call flow for audit. Compliance, legal and your regulator have all seen the diagram. "The model decides" is not an acceptable replacement; you owe an equivalent artefact — the policy the agent is instructed with, versioned, plus a per-call trace showing which path was taken.
- Guaranteed script delivery. Mandatory disclosures fired every time because a node played them. A model can skip, paraphrase or reorder one. Play required disclosures from deterministic code outside the model's control, not from a system prompt asking nicely.
- Predictable cost per call. An IVR minute cost the same on every call. An agent minute varies with tokens, tool calls and how long the caller talks, and the long tail is where the money goes — the same shape as forecasting agent spend, with a hard per-call ceiling that hands off rather than continuing.
- A stable evidentiary record. Recording consent, retention and redaction now apply to a transcript, a model context, tool arguments and a trace as well as the audio — five artefacts where you had one, as recording, consent and redaction sets out.
- No prompt-injection surface. The caller could press digits. Now they can speak arbitrary text into a context that reaches a model with tool access; a caller reading out an "instruction" is the voice channel's version of the same untrusted-input problem.
Write these five down as a deliverable list at kickoff, owned by name. Every one of them is a thing the IVR provided for free that your programme is now on the hook for, and each is discovered late — usually by compliance, usually the week before launch.
Run the first eight weeks on real calls, not on a script.
The evaluation set that matters is built from the audio you already have. Everything else is a rehearsal for a call nobody made.
- Golden set from recordings, not from written test cases. Sample real calls per intent, including the bad audio, the code-switching, the background noise and the caller who explains their problem in the wrong order. Replay them; a script-derived suite passes at 95% and predicts nothing.
- Score the transfer as well as the resolution. A clean, early handoff is a good outcome. Grade escalation precision and recall alongside resolution, per escalation and warm transfer, or you will optimise the agent into clinging to calls it should have passed on.
- Listen to fifty calls a week, by name. Somebody senior, every week, for the first quarter. Every team that skipped this discovered their worst failure mode from a customer complaint instead of from their own queue.
- Watch the human side of the seam. Agents receiving transfers will tell you within days which intents are failing and how, and they are the cheapest evaluation instrument you own.
Start with one intent, on one number, with the legacy IVR as the fallback and a rollback that is a config flip — and pick the metric before the pilot: resolution without a callback in 72 hours, with abandonment inside the agent as the guardrail that can veto a launch. Cluster the zero-out transcripts to choose that first intent, keep DTMF for digits, and put mandatory disclosures in deterministic code. Teams that do this ship a narrow thing that works in six weeks; teams that rebuild the menu tree in natural language spend two quarters producing a more expensive IVR that is harder to audit. For the surrounding operating model, see customer support agents.