Playbooks / Voice & Realtime Agents
Voice & Realtime Agents
Realtime voice agents — speech stack, turn-taking, barge-in, latency budgets, voice-specific tooling and state.
- Realtime Agent ArchitectureCascade (STT→LLM→TTS) vs native speech-to-speech, the stateful audio transport, and the one decision everything else hangs on: where the agent loop lives.
- The Latency BudgetThe sub-second turn accounted for line by line: where the milliseconds go, why endpointing is the biggest slice, and perceived vs actual latency.
- Turn-Taking & Barge-InVAD vs endpointing, semantic end-of-turn detection, mandatory barge-in, echo cancellation as a prerequisite, and backchannels vs real interruptions.
- STT, TTS & Speech-to-SpeechStreaming STT, the transcription-error tax, TTS time-to-first-audio, native audio models, and why 8 kHz telephony changes every benchmark.
- Tool Use & State in VoiceCalling tools without dead air: preambles, async/parallel tool runs, confirm-by-ear before mutating, and slot state across an interruptible call.
- Voice Agent Failure ModesHallucinated hearing, dead air, the infinite apology loop, the latency death spiral, and the escalation/handoff you must design for.
- Outbound voice agentsAgents that **make** the call instead of answering it — pacing, abandonment, identity disclosure, and the regulatory landmines that turn a clever demo into a fine.
- Evaluating Voice AgentsTranscript evals score the one layer that was not broken: build the golden set from recorded audio, measure entity error rate rather than WER, and treat timing as a first-class score.
- Telephony & PSTN IntegrationHalf the turn latency, all the audio quality and whether the call connects at all live in a carrier path you cannot profile: the fixed transport tax, your number as a reputation asset, the missing metadata channel, and consent as a code path rather than a prompt.
- Multilingual & Code-Switching Voice AgentsAdding a language breaks recognition, not generation — and per-turn language lock, the standard fix, is exactly what a bilingual caller violates in their first sentence, with the damage landing on the names and numbers you were about to pass to a tool.
- Caller Authentication for Voice AgentsThree seconds of audio clones a customer and roughly one in five biometric fraud attempts is now a deepfake, so a voiceprint identifies but no longer authenticates — move the proof out of the audio channel, bind it to the action rather than the call, and notice that your own outbound agent is normalising the attack.
- Retrieval Inside the Voice TurnA grounded answer has to start leaving the speaker about 800ms after the caller stops, and a retrieve-rewrite-rerank chain spends most of that before the model sees a document — so retrieval latency is an accuracy metric: start querying on the partial transcript, precompute the head of the question distribution, and delete the stages you were told were mandatory.
- Escalation & Warm TransferA voice agent can nail every turn and still lose the customer in the handoff, because the human answers as if the call never happened — so the handoff is judged by how much the caller repeats, not by whether it connected: escalate before the caller asks, carry a context packet (verified identity, the caller’s own words, actions already taken) that lands before the human’s first word, always design the no-human branch, and page yourself on repeat rate rather than transfer rate.
- After-Call Work & CRM WritebackThe caller hears the conversation once; the disposition code and summary are read for years by routers, analysts and auditors — and your voice eval stops when the caller hangs up. The disposition is a classification problem wearing a generation problem’s clothes, your wrap-up taxonomy is probably already broken, and the summary must quote rather than characterise, because ASR errors concentrate exactly on the names and amounts a durable record cannot get wrong.
- Recording, Consent & RedactionYour retention rule points at the call recording, and the recording is now the least interesting copy: the same minute also lives as a transcript, a model context, tool arguments, a trace span and a CRM summary — five artefacts the agent created that inherited no policy. Gate the buffer on consent rather than call setup, keep card data off the agent leg entirely because a model cannot look away, and run one versioned redactor in front of every sink at write time.
- Replacing an IVRThe two artefacts you start from are both traps: the menu tree records what touch-tone could express rather than what callers want, and containment — quoted at 5–10% for legacy IVRs against 60–90% for voice agents — scores the caller who gave up as a success. Build the intent inventory from the zero-out transcripts, migrate one intent at a time in front of the IVR you already trust so rollback is a config flip, measure resolution without a callback in 72 hours with in-agent abandonment as the veto, and keep DTMF for digits.
- Accessible Voice AgentsYour agent does not have a failure rate, it has one per kind of voice, and the cohorts that fail hardest — disordered or slow speech, strong accents, older callers, anyone on a relay service — have the fewest alternatives, so their failures arrive as hang-ups and never reach your metrics. Silence-based endpointing is not merely blind to them, it is the mechanism: stratify by speech rate rather than by any label, raise the threshold permanently after the first re-prompt, keep DTMF and a human path live at every turn, and gate releases on worst-cohort over median-cohort success.
- Alphanumerics Over VoiceYour transcription is excellent and your order lookups fail, because a booking reference has no language model behind it and exact match is per-character accuracy raised to the length — 97% per character is 73.7% on ten characters. Provider spelling modes and keyword biasing buy percentage points; what changes the shape is resolving against the small candidate set the caller's phone number already gives you, chunked capture with per-chunk readback when you truly must, confirmation set by consequence, and taking the finding upstream to the identifier format.
- Disclosing the Agent on a CallThe one sentence you are legally required to say is the one most likely to be cut by your own barge-in, and it logs as played. Three obligations with different shapes — EU AI Act Art. 50 up front since 2 August 2026, Utah on request, Utah again prominently for high-risk work — need three code paths: a short non-interruptible opener logged on final frame, a must_disclose predicate re-evaluated at transfer and party-join and resume, and a deterministic "are you a bot" intent with a constant-string answer, because a system-prompt line makes statutory compliance a model behaviour.
- Card Payments Over VoiceEvery contact-centre descoping technique works by removing a listener from the audio path, and a voice agent is not a listener — it is the call, so one spoken card number lands in six artefacts at once, two of which PCI DSS says may never be stored post-authorisation. Three architectures keep the data out of the model and they differ only in what they cost the caller. The change to make today is the tool signature: if a PAN can be an argument, your trace store is in scope.