Playbooks / Voice & Realtime Agents

Voice & Realtime Agents

Realtime voice agents — speech stack, turn-taking, barge-in, latency budgets, voice-specific tooling and state.

  1. Realtime Agent Architecture
    Cascade (STT→LLM→TTS) vs native speech-to-speech, the stateful audio transport, and the one decision everything else hangs on: where the agent loop lives.
  2. The Latency Budget
    The sub-second turn accounted for line by line: where the milliseconds go, why endpointing is the biggest slice, and perceived vs actual latency.
  3. Turn-Taking & Barge-In
    VAD vs endpointing, semantic end-of-turn detection, mandatory barge-in, echo cancellation as a prerequisite, and backchannels vs real interruptions.
  4. STT, TTS & Speech-to-Speech
    Streaming STT, the transcription-error tax, TTS time-to-first-audio, native audio models, and why 8 kHz telephony changes every benchmark.
  5. Tool Use & State in Voice
    Calling tools without dead air: preambles, async/parallel tool runs, confirm-by-ear before mutating, and slot state across an interruptible call.
  6. Voice Agent Failure Modes
    Hallucinated hearing, dead air, the infinite apology loop, the latency death spiral, and the escalation/handoff you must design for.
  7. Outbound voice agents
    Agents that **make** the call instead of answering it — pacing, abandonment, identity disclosure, and the regulatory landmines that turn a clever demo into a fine.
  8. Evaluating Voice Agents
    Transcript evals score the one layer that was not broken: build the golden set from recorded audio, measure entity error rate rather than WER, and treat timing as a first-class score.
  9. Telephony & PSTN Integration
    Half the turn latency, all the audio quality and whether the call connects at all live in a carrier path you cannot profile: the fixed transport tax, your number as a reputation asset, the missing metadata channel, and consent as a code path rather than a prompt.
  10. Multilingual & Code-Switching Voice Agents
    Adding a language breaks recognition, not generation — and per-turn language lock, the standard fix, is exactly what a bilingual caller violates in their first sentence, with the damage landing on the names and numbers you were about to pass to a tool.
  11. Caller Authentication for Voice Agents
    Three seconds of audio clones a customer and roughly one in five biometric fraud attempts is now a deepfake, so a voiceprint identifies but no longer authenticates — move the proof out of the audio channel, bind it to the action rather than the call, and notice that your own outbound agent is normalising the attack.
  12. Retrieval Inside the Voice Turn
    A grounded answer has to start leaving the speaker about 800ms after the caller stops, and a retrieve-rewrite-rerank chain spends most of that before the model sees a document — so retrieval latency is an accuracy metric: start querying on the partial transcript, precompute the head of the question distribution, and delete the stages you were told were mandatory.
  13. Escalation & Warm Transfer
    A voice agent can nail every turn and still lose the customer in the handoff, because the human answers as if the call never happened — so the handoff is judged by how much the caller repeats, not by whether it connected: escalate before the caller asks, carry a context packet (verified identity, the caller’s own words, actions already taken) that lands before the human’s first word, always design the no-human branch, and page yourself on repeat rate rather than transfer rate.
  14. After-Call Work & CRM Writeback
    The caller hears the conversation once; the disposition code and summary are read for years by routers, analysts and auditors — and your voice eval stops when the caller hangs up. The disposition is a classification problem wearing a generation problem’s clothes, your wrap-up taxonomy is probably already broken, and the summary must quote rather than characterise, because ASR errors concentrate exactly on the names and amounts a durable record cannot get wrong.
  15. Recording, Consent & Redaction
    Your retention rule points at the call recording, and the recording is now the least interesting copy: the same minute also lives as a transcript, a model context, tool arguments, a trace span and a CRM summary — five artefacts the agent created that inherited no policy. Gate the buffer on consent rather than call setup, keep card data off the agent leg entirely because a model cannot look away, and run one versioned redactor in front of every sink at write time.
  16. Replacing an IVR
    The two artefacts you start from are both traps: the menu tree records what touch-tone could express rather than what callers want, and containment — quoted at 5–10% for legacy IVRs against 60–90% for voice agents — scores the caller who gave up as a success. Build the intent inventory from the zero-out transcripts, migrate one intent at a time in front of the IVR you already trust so rollback is a config flip, measure resolution without a callback in 72 hours with in-agent abandonment as the veto, and keep DTMF for digits.
  17. Accessible Voice Agents
    Your agent does not have a failure rate, it has one per kind of voice, and the cohorts that fail hardest — disordered or slow speech, strong accents, older callers, anyone on a relay service — have the fewest alternatives, so their failures arrive as hang-ups and never reach your metrics. Silence-based endpointing is not merely blind to them, it is the mechanism: stratify by speech rate rather than by any label, raise the threshold permanently after the first re-prompt, keep DTMF and a human path live at every turn, and gate releases on worst-cohort over median-cohort success.
  18. Alphanumerics Over Voice
    Your transcription is excellent and your order lookups fail, because a booking reference has no language model behind it and exact match is per-character accuracy raised to the length — 97% per character is 73.7% on ten characters. Provider spelling modes and keyword biasing buy percentage points; what changes the shape is resolving against the small candidate set the caller's phone number already gives you, chunked capture with per-chunk readback when you truly must, confirmation set by consequence, and taking the finding upstream to the identifier format.
  19. Disclosing the Agent on a Call
    The one sentence you are legally required to say is the one most likely to be cut by your own barge-in, and it logs as played. Three obligations with different shapes — EU AI Act Art. 50 up front since 2 August 2026, Utah on request, Utah again prominently for high-risk work — need three code paths: a short non-interruptible opener logged on final frame, a must_disclose predicate re-evaluated at transfer and party-join and resume, and a deterministic "are you a bot" intent with a constant-string answer, because a system-prompt line makes statutory compliance a model behaviour.
  20. Card Payments Over Voice
    Every contact-centre descoping technique works by removing a listener from the audio path, and a voice agent is not a listener — it is the call, so one spoken card number lands in six artefacts at once, two of which PCI DSS says may never be stored post-authorisation. Three architectures keep the data out of the model and they differ only in what they cost the caller. The change to make today is the tool signature: if a PAN can be an argument, your trace store is in scope.