AI Blog

Full duplex deletes the turn — and the turn was your commit point

GPT-Live-1 landed in the API on 10 September and listens while it speaks, which reads as a naturalness upgrade and is actually a schema change. End-of-turn was the event your voice agent used to decide when to call a tool, when to write a log line, when to run a guardrail and when to stop the meter — and a full-duplex model never fires it. The fix is not a better threshold; it is naming your own commit points and pricing a meter that now runs on wall clock instead of speech.

By Agentic AI Wiki 14 min read

The voice model OpenAI put in the API on 10 September keeps listening while it is still talking, and the interesting consequence is not that conversations feel better. It is that end-of-turn — the event your stack used to decide when to call a tool, when to write a trace, when to run a guardrail and when to stop the meter — no longer happens. Every one of those decisions was hanging off a signal that a full-duplex model does not emit.

At a glance

GPT-Live-1 is deliberately not a brain. It is the mouth and ears, sold separately from the thinking.

PropertyGPT-Live-1What it replaces
Conversation model Full duplex — listens and speaks simultaneously. Half duplex: detect end of speech, then respond.
Reasoning and tools Delegated to a backend text model you choose. Handled inside the realtime model.
Price $0.05/minute, billed by the second — voice layer only. Audio metered as tokens.
Transports WebRTC, WebSocket and SIP, via v1/live/sessions. Same three, different session object.
Where the end-of-turn event went Two panels. On the left, a half-duplex loop: caller audio feeds voice activity detection, which feeds endpointing, which emits a single end-of-turn event drawn as a solid accent box. Four arrows leave that event and reach tool dispatch, trace write, guardrail check and billing tick. On the right, a full-duplex loop: caller audio and agent audio flow together through one model with no endpointing stage and no end-of-turn event, and the same four consumers are drawn detached with question marks, waiting for a signal that never arrives. A band across the bottom names three commit points you declare instead: a filled and stable slot fires the tool, a backend model invocation opens a span, and the agent starting or stopping audio marks the timeline. One event, four subscribers Half duplex Caller audio Voice activity detection Endpointing end-of-turn Tool dispatch Trace write Guardrail check Billing tick Full duplex Caller audio Agent audio One model continuous, both ways no end-of-turn Tool dispatch ? Trace write ? Guardrail check ? Billing tick ? What you declare instead Slot filled and stable fire the tool Backend model invoked open a span Audio starts or stops mark the timeline
Four subsystems were subscribed to one event. Removing the event does not remove the subscribers.

What actually shipped

On 10 September 2026 OpenAI made gpt-live-1 generally available in the API at five cents a minute, billed per second. The model handles incoming and outgoing audio together: it can react to an interruption, a backchannel "mm-hm", a mid-sentence change of direction or background speech while its own audio is still playing.

The architectural choice underneath is the one worth reading twice. GPT-Live-1 does not do the thinking. You pair it with a backend text model — OpenAI names GPT-6 Astra, Luna and Terra, and third-party models are allowed — and that model does the reasoning and the tool calls while the voice layer keeps the conversation moving. The reported numbers come from that pairing rather than from the voice model alone: on Full Duplex Bench it scores 80.1% on interactivity against 45.4% for GPT-Realtime-2.1, with turn-taking latency down to 0.8 s from 1.4 s; on Full Duplex Bench v3, run with a Terra backend at low reasoning effort, tool-calling Pass@1 is 87.0% against 60.0% and 58.0% for the comparison configurations. Paired with GPT-6 Astra at medium reasoning effort, OpenAI reports a first-place result on Tau3, an end-to-end benchmark of spoken customer-service tasks.

Operationally it is still a beta-shaped product. Concurrent sessions cap at 25 on tier 1 and 500 on tier 5, and the free tier is not supported — which means the first capacity plan you write is a tier-upgrade request, not a load test.

Note what the benchmark configuration implies: every headline number is a tuple of voice model, backend model and reasoning effort. "GPT-Live-1 scores 87%" is not a statement about a model. Record all three when you quote it, and all three when you measure your own.

The turn was load-bearing, and almost nobody wrote that down

Where the commit point lives across three voice architectures Three columns. Cascaded half duplex: endpointing runs in your own code and emits end-of-turn, so the commit point is a silence threshold you can tune and revert. Speech-to-speech half duplex: endpointing has moved inside the model but a turn boundary is still exposed as an event you subscribe to. Full duplex: there is no turn boundary at all, so the commit point becomes a policy you write and must maintain. Where the commit point lives Cascaded half duplex Endpointing is your code. You emit end-of-turn and decide everything from it. Speech-to-speech Endpointing moved inside the model, but a turn is still exposed as an event. Full duplex No turn boundary exists. Nothing fires when the caller stops speaking. Commit point Commit point Commit point A silence threshold you tune, diff and revert. A model event you subscribe to, not tune. A policy you write, own and maintain.
Two architecture changes moved the turn boundary. The third removes it.

In a cascaded stack, endpointing is a decision your code makes: voice activity detection says whether there is speech in this frame, and an endpointer decides whether the caller is done. That second decision is the most consequential line in a voice agent, and teams tune it for months. What gets noticed less is how much else was quietly bolted to it.

  • Tool dispatch. "Call the tool when the caller finishes" is the default design, because a tool call mid-utterance risks acting on half a sentence. End-of-turn was the guard.
  • Trace boundaries. A span per turn is how voice traces are shaped. Without turns, a thirty-minute call is one span with no internal structure, and the trace stops being readable at exactly the length where you need it.
  • Guardrails. An output check that runs on the completed utterance cannot run on an utterance that is still being emitted while the caller talks over it.
  • Evaluation. Turn-level scoring — was this response correct, given this user turn — is how nearly every voice eval suite is built. The voice-eval vendors segment by turn because that is the unit that existed.
  • Human handoff. "Escalate at the end of the current turn" is the polite version of a transfer. There is no end of the current turn.

None of these are hard problems once you see them. They are dangerous because they are invisible: nothing in your code says on_end_of_turn for all five. The dependency is distributed across a state machine, a tracing decorator, a guardrail middleware and a test harness, and the first three will keep running without error. They will simply run at the wrong moments, or never.

The failure this produces is not a crash. It is a tool called against a sentence the caller was still amending, a guardrail that passes because it evaluated an empty buffer, and a trace that says one thing happened. Every one of those looks like a model regression in your dashboard.

What replaces the turn: commit points you declare

The honest answer is that the turn was always a proxy for something else — the moment at which the agent's understanding is stable enough to act on. Half duplex let you conflate the two, because the caller stopping talking was a decent approximation of the caller being finished. Full duplex forces the distinction into the open, which is uncomfortable and also correct.

So name the commit points instead of inheriting them:

  • Commit on a completed slot, not on silence. A booking agent does not need the caller to stop; it needs a date, a party size and a name. Fire the tool when the slot is filled and stable for a short window, and let the caller keep talking. This is what explicit dialogue state was for, and full duplex makes it mandatory rather than tidy.
  • Make retraction a first-class path. If you act before the caller is done, they will sometimes correct you — "no, Tuesday" — after the tool ran. That is now the ordinary case, not an edge case, so every commit-early tool needs a compensating action. Read-only lookups are free to fire early; anything that mutates needs the same idempotency key and reconciliation you would put on a retried booking.
  • Span on activity, not on turns. Emit a span when a tool fires, when the backend model is invoked, when the agent starts and stops speaking. The call becomes a timeline of overlapping intervals, which is what it always was.
  • Move guardrails to the stream. An output check that needs a finished utterance is the wrong shape now. Check the backend model's tool arguments — those are discrete — and accept that the audio itself is governed by the model's own behaviour.

The uncomfortable part: the decision of when to speak has moved inside the model's weights. In a cascade you tuned a number, and a number is reviewable, diffable and revertible. A learned interruption policy is none of those. It is the same trade the industry made when it moved from rule-based dialogue to LLMs, arriving one layer lower, and it deserves the same response — pin the version, keep a frozen scenario set, and re-baseline on every change.

Two meters, opposite shapes

Reported Full Duplex Bench results Horizontal bar chart with two groups. Full Duplex Bench interactivity: GPT-Live-1 scores 80.1 percent and GPT-Realtime-2.1 scores 45.4 percent. Full Duplex Bench version three tool-calling Pass at 1, run with a Terra backend at low reasoning effort: GPT-Live-1 scores 87.0 percent against comparison configurations at 60.0 and 58.0 percent. Reported score (%) Full Duplex Bench — interactivity GPT-Live-1 80.1 GPT-Realtime-2.1 45.4 Full Duplex Bench v3 — tool-calling Pass@1 (Terra backend, low effort) GPT-Live-1 87.0 Comparison config A 60.0 Comparison config B 58.0 25 50 75 100 Each bar is a configuration — voice layer, backend model and reasoning effort together. Turn-taking latency, not to scale: 0.8 s for GPT-Live-1 against 1.4 s for GPT-Realtime-2.1.
Every bar is a configuration, not a model: voice layer, backend model and reasoning effort together.

Unbundling the voice layer from the thinking splits your bill into two meters that behave nothing alike, and the split is easy to miss because only one of them appears on the launch page.

The voice layer is five cents a minute of wall clock. It runs while the caller talks, while the agent talks, and — this is the part worth internalising — while nobody talks. A caller hunting for their account number generates no tokens and costs the same as a caller mid-sentence. Token-metered audio priced what was said; a per-minute layer prices how long the line was open. Anything that lengthens a call now costs money at a flat rate, which quietly re-prices the whole design: a slow backend model, a retrieval step, a hold while a tool runs, a caller who rambles.

The backend model is the opposite: metered in tokens, indifferent to duration, and driven by how much you make it think. Reasoning effort is a dial that moves your cost per call without moving a single minute of audio — and the benchmark table above shows it also moves quality, so the two meters are coupled through a knob that does not appear on either invoice.

Which makes the number to track obvious and slightly unusual: cost per resolved call, decomposed into minutes and tokens. Not cost per minute, not cost per token. A change that cuts backend spend by shifting to a cheaper model but adds eight seconds of thinking time per exchange can easily lose money, and neither meter alone will tell you. This is the same discipline agent unit economics asks for, applied to a system where latency has become a direct cost line rather than a UX concern.

A concrete audit: take a week of call recordings, measure total silence, and multiply by $0.05/minute. That number is what your endpointing patience used to be free and now is not. If it is large, the fix is usually not a faster model — it is a shorter conversation design.

Who should move, and when

SituationAdopt full duplex nowStay half duplex
Conversational, interruption-heavy, low stakes per utterance Yes — this is what it is for. Only if your barge-in already works well.
Every action is a mutation (payments, bookings, dispatch) Only with compensating actions built first. Yes — the turn boundary is cheap safety.
Regulated call flows with scripted disclosures Not yet — a disclosure the caller talked over is a finding. Yes.
Your eval suite scores per turn Rebuild the eval first, then adopt. Yes, until then.
Long silences are structural (lookups, hold music, IVR legacy) Model the minute cost before committing. Possibly cheaper.

The broader read is that OpenAI has done to the voice stack what the Agents API did to the control loop: taken a layer teams were building badly, made it excellent, and moved a decision you used to own inside a version you cannot pin. That is a good trade in both cases, and in both cases the price is your ability to attribute a regression. The defence is identical too — one frozen scenario you re-run daily, and a number you compute yourself.

FAQ

Does full duplex mean I no longer need voice activity detection?

You no longer need it to decide when the agent should reply — the model does that. You still want it for the things VAD was separately good at: detecting that a call has gone silent, measuring talk-over rates, and segmenting recordings for review.

Why does GPT-Live-1 need a separate backend model?

Because keeping a conversation flowing and reasoning carefully are different jobs with different latency budgets. Splitting them lets the voice layer stay responsive while a slower model thinks, and lets you choose how hard that model thinks per call. It also means every benchmark number is a property of the pair, not of either half.

Is $0.05 per minute cheaper than what I pay today?

It is not comparable on its own, because it buys only the voice layer. Add the backend model's tokens and any tools. The structural change is that the voice portion is now billed on elapsed time rather than on audio tokens, so silence and slow tool calls have a price they did not have before.

What breaks first when I switch an existing agent over?

Usually the tool-dispatch guard and the eval suite, in that order. Anything with the shape "wait for the user to finish, then act" has no trigger, and any scoring harness segmented by turn has no segments.

Can I keep a human-escalation path in a full-duplex call?

Yes, but the handoff point has to be declared rather than inferred. Pick an explicit condition — a confidence threshold, a caller phrase, a failed tool — and have the agent stop speaking on it, instead of waiting for a turn to end. See escalation and warm transfer.

Further reading

On this wiki:

Sources: