Full-duplex speech.
A voice agent that can talk while it listens does not just feel more natural — it stops producing the one event the rest of your system was built around. Half-duplex agents wait for silence, decide the human is finished, and act; full-duplex models never make that decision, so the tuning knob you used to own becomes a behaviour you can only elicit. Treat full duplex as a change of unit, not a change of speed.
The turn is an artefact of the pipeline, not a property of speech.
Human conversation is not strictly turn-taking. People overlap constantly: backchannels ("mm-hm", "right") land on top of the speaker without claiming the floor, listeners latch onto the end of a phrase before it finishes, and a correction arrives mid-sentence precisely because waiting would be rude. Linguists have measured gaps between turns at around 200 ms on average — far shorter than the time it takes to plan a sentence, which means listeners start planning before the speaker stops.
Machine voice pipelines could not do any of that, for a mechanical reason: they process one direction at a time. Audio comes in, a boundary is declared, a response is generated, audio goes out. Half duplex is that pattern, and it is borrowed from radio — one party transmits while the other receives. The "turn" in a voice agent is where that switch happens, and it exists because the architecture needed a switch, not because conversation has one.
- Half duplex. One direction at a time. Requires a decision — is the human finished? — that a machine must make explicitly and a human never makes at all.
- Barge-in. The patch: detect that the human started talking, stop the outbound audio, discard the rest. It gives the appearance of overlap while remaining half duplex underneath, which is why interrupting a voice agent so often feels like hanging up on it.
- Full duplex. Both directions carried at once, continuously, with no switch. There is nothing to detect and nothing to stop, because the agent was never blocked on the human finishing.
The tell that you are still half duplex: somewhere in the code there is a timeout in milliseconds. Endpointing is the art of choosing that number, and the reason it is an art is that no single value is both responsive and patient. Full duplex does not give you a better number — it removes the question.
What full duplex actually requires.
"Listen and speak at the same time" is easy to say and expensive to build, because two things become simultaneously true that a half-duplex system never had to reconcile.
- The model hears itself. If the agent's audio is playing while the microphone is open, the agent's own voice is in the input stream. Acoustic echo cancellation stops being an audio-quality nicety and becomes a correctness requirement: without it, the model transcribes its own sentence as the user's reply and answers it. This is why full duplex is much harder over a speakerphone or a drive-thru lane than over a headset.
- The model must decide whether to speak, continuously. In a half-duplex loop, "should I talk now?" is answered for the model by the endpointer. In a full-duplex loop it is a decision the model makes on every frame, alongside what to say. That decision has no name in the half-duplex vocabulary because nothing made it.
- Overlap has to mean something. A user talking over the agent might be interrupting, agreeing, or thinking aloud. A model that stops on every overlap is worse than a half-duplex one; a model that ignores overlap is a bulldozer. Handling this well is the actual product, and it is learned, not configured.
Because those are hard and latency-critical, the shipping architecture usually splits the job: a small, fast full-duplex model owns the audio and the decision to speak, and a separate, slower text model does the reasoning and the tool calls behind it. That split is worth understanding before you read any benchmark, because a reported score belongs to the pair — voice model, backend model, and how hard the backend was told to think — and not to either half alone.
A threshold you tuned becomes a behaviour you elicit.
This is the trade, and it is the same one the field has already made twice at higher layers. Rule-based dialogue became LLM dialogue; hand-written tool selection became model tool selection. Each time, a thing you could read, diff and revert turned into a thing you can only measure.
An endpointing threshold is 800 in a config file. You can change it, review the change, correlate it with a metric, and roll it back on a Friday. A full-duplex model's interruption policy is distributed across its weights. You can prompt around the edges, you can pick a different model, and that is the extent of your control.
- Regressions stop being attributable. When callers start reporting that the agent talks over them, there is no line to blame. The candidates are the model version, the prompt, the audio path and the caller population — which is why pinning versions and keeping a frozen scenario set matters more here, not less.
- Per-turn evaluation stops working. Nearly every voice eval suite scores "given this user turn, was this agent turn correct?" With no turns, there are no segments. Scoring moves to intervals and outcomes: did the agent talk over the caller, how long was the caller waiting, was the task completed. See evaluating voice agents.
- Acting early becomes ordinary. Without an end-of-turn guard, an agent can commit to a tool call while the caller is still amending the request. Read-only lookups are fine; anything that mutates needs a compensating action, because "no, Tuesday" arriving after the booking is now the normal case rather than an edge case.
The failure mode is quiet. Nothing throws. A tool fires against half a sentence, a guardrail evaluates an empty buffer, a trace records one event where two overlapped — and all three surface in your dashboard as a model regression rather than as a missing boundary.
The meter changes shape too.
Half-duplex voice was usually metered in audio tokens, which priced what was said. Full-duplex voice layers are priced per minute of elapsed time, which prices how long the line was open — and those are not the same bill at all.
Silence is the clearest example. A caller hunting for an account number produces no tokens and, under per-minute pricing, costs exactly as much as a caller mid-sentence. So does a pause while a tool runs, a slow backend model, and a conversation design that asks three questions where one would do. Latency stops being only a user-experience concern and becomes a direct cost line, which changes which optimisations are worth making: shortening the conversation beats speeding up the model more often than it used to.
Meanwhile the backend model is still metered in tokens and is indifferent to duration. You now have two meters with opposite shapes, coupled by a reasoning-effort dial that appears on neither invoice. The only number that reconciles them is cost per completed task, decomposed into minutes and tokens — the discipline agent cost control asks for, applied to a system where wall-clock time has a price.
If you do one thing before adopting full duplex: write down, explicitly, what your agent's commit points are — the conditions under which it may call a tool, open a trace span, run a check, or hand off to a human. In a half-duplex system you inherited all four from end-of-turn without ever deciding them. Full duplex does not break those subsystems; it removes the signal they were silently subscribed to, and the only ones that keep working are the ones you have named. Related: voice and realtime agents for the cascade-versus-speech-to-speech choice underneath this, the latency budget for where the milliseconds go, and human-in-the-loop for placing the checkpoint once the turn can no longer hold it.