The voice model OpenAI put in the API on 10 September keeps listening while it is still talking, and the interesting consequence is not that conversations feel better. It is that end-of-turn — the event your stack used to decide when to call a tool, when to write a trace, when to run a guardrail and when to stop the meter — no longer happens. Every one of those decisions was hanging off a signal that a full-duplex model does not emit.
At a glance
GPT-Live-1 is deliberately not a brain. It is the mouth and ears, sold separately from the thinking.
| Property | GPT-Live-1 | What it replaces |
|---|---|---|
| Conversation model | Full duplex — listens and speaks simultaneously. | Half duplex: detect end of speech, then respond. |
| Reasoning and tools | Delegated to a backend text model you choose. | Handled inside the realtime model. |
| Price | $0.05/minute, billed by the second — voice layer only. | Audio metered as tokens. |
| Transports | WebRTC, WebSocket and SIP, via v1/live/sessions. |
Same three, different session object. |
What actually shipped
On 10 September 2026 OpenAI made gpt-live-1 generally available in the API at five cents a minute, billed per second. The model handles incoming and outgoing audio together: it can react to an interruption, a backchannel "mm-hm", a mid-sentence change of direction or background speech while its own audio is still playing.
The architectural choice underneath is the one worth reading twice. GPT-Live-1 does not do the thinking. You pair it with a backend text model — OpenAI names GPT-6 Astra, Luna and Terra, and third-party models are allowed — and that model does the reasoning and the tool calls while the voice layer keeps the conversation moving. The reported numbers come from that pairing rather than from the voice model alone: on Full Duplex Bench it scores 80.1% on interactivity against 45.4% for GPT-Realtime-2.1, with turn-taking latency down to 0.8 s from 1.4 s; on Full Duplex Bench v3, run with a Terra backend at low reasoning effort, tool-calling Pass@1 is 87.0% against 60.0% and 58.0% for the comparison configurations. Paired with GPT-6 Astra at medium reasoning effort, OpenAI reports a first-place result on Tau3, an end-to-end benchmark of spoken customer-service tasks.
Operationally it is still a beta-shaped product. Concurrent sessions cap at 25 on tier 1 and 500 on tier 5, and the free tier is not supported — which means the first capacity plan you write is a tier-upgrade request, not a load test.
Note what the benchmark configuration implies: every headline number is a tuple of voice model, backend model and reasoning effort. "GPT-Live-1 scores 87%" is not a statement about a model. Record all three when you quote it, and all three when you measure your own.
The turn was load-bearing, and almost nobody wrote that down
In a cascaded stack, endpointing is a decision your code makes: voice activity detection says whether there is speech in this frame, and an endpointer decides whether the caller is done. That second decision is the most consequential line in a voice agent, and teams tune it for months. What gets noticed less is how much else was quietly bolted to it.
- Tool dispatch. "Call the tool when the caller finishes" is the default design, because a tool call mid-utterance risks acting on half a sentence. End-of-turn was the guard.
- Trace boundaries. A span per turn is how voice traces are shaped. Without turns, a thirty-minute call is one span with no internal structure, and the trace stops being readable at exactly the length where you need it.
- Guardrails. An output check that runs on the completed utterance cannot run on an utterance that is still being emitted while the caller talks over it.
- Evaluation. Turn-level scoring — was this response correct, given this user turn — is how nearly every voice eval suite is built. The voice-eval vendors segment by turn because that is the unit that existed.
- Human handoff. "Escalate at the end of the current turn" is the polite version of a transfer. There is no end of the current turn.
None of these are hard problems once you see them. They are dangerous because they are invisible: nothing in your code says on_end_of_turn for all five. The dependency is distributed across a state machine, a tracing decorator, a guardrail middleware and a test harness, and the first three will keep running without error. They will simply run at the wrong moments, or never.
The failure this produces is not a crash. It is a tool called against a sentence the caller was still amending, a guardrail that passes because it evaluated an empty buffer, and a trace that says one thing happened. Every one of those looks like a model regression in your dashboard.
What replaces the turn: commit points you declare
The honest answer is that the turn was always a proxy for something else — the moment at which the agent's understanding is stable enough to act on. Half duplex let you conflate the two, because the caller stopping talking was a decent approximation of the caller being finished. Full duplex forces the distinction into the open, which is uncomfortable and also correct.
So name the commit points instead of inheriting them:
- Commit on a completed slot, not on silence. A booking agent does not need the caller to stop; it needs a date, a party size and a name. Fire the tool when the slot is filled and stable for a short window, and let the caller keep talking. This is what explicit dialogue state was for, and full duplex makes it mandatory rather than tidy.
- Make retraction a first-class path. If you act before the caller is done, they will sometimes correct you — "no, Tuesday" — after the tool ran. That is now the ordinary case, not an edge case, so every commit-early tool needs a compensating action. Read-only lookups are free to fire early; anything that mutates needs the same idempotency key and reconciliation you would put on a retried booking.
- Span on activity, not on turns. Emit a span when a tool fires, when the backend model is invoked, when the agent starts and stops speaking. The call becomes a timeline of overlapping intervals, which is what it always was.
- Move guardrails to the stream. An output check that needs a finished utterance is the wrong shape now. Check the backend model's tool arguments — those are discrete — and accept that the audio itself is governed by the model's own behaviour.
The uncomfortable part: the decision of when to speak has moved inside the model's weights. In a cascade you tuned a number, and a number is reviewable, diffable and revertible. A learned interruption policy is none of those. It is the same trade the industry made when it moved from rule-based dialogue to LLMs, arriving one layer lower, and it deserves the same response — pin the version, keep a frozen scenario set, and re-baseline on every change.
Two meters, opposite shapes
Unbundling the voice layer from the thinking splits your bill into two meters that behave nothing alike, and the split is easy to miss because only one of them appears on the launch page.
The voice layer is five cents a minute of wall clock. It runs while the caller talks, while the agent talks, and — this is the part worth internalising — while nobody talks. A caller hunting for their account number generates no tokens and costs the same as a caller mid-sentence. Token-metered audio priced what was said; a per-minute layer prices how long the line was open. Anything that lengthens a call now costs money at a flat rate, which quietly re-prices the whole design: a slow backend model, a retrieval step, a hold while a tool runs, a caller who rambles.
The backend model is the opposite: metered in tokens, indifferent to duration, and driven by how much you make it think. Reasoning effort is a dial that moves your cost per call without moving a single minute of audio — and the benchmark table above shows it also moves quality, so the two meters are coupled through a knob that does not appear on either invoice.
Which makes the number to track obvious and slightly unusual: cost per resolved call, decomposed into minutes and tokens. Not cost per minute, not cost per token. A change that cuts backend spend by shifting to a cheaper model but adds eight seconds of thinking time per exchange can easily lose money, and neither meter alone will tell you. This is the same discipline agent unit economics asks for, applied to a system where latency has become a direct cost line rather than a UX concern.
A concrete audit: take a week of call recordings, measure total silence, and multiply by $0.05/minute. That number is what your endpointing patience used to be free and now is not. If it is large, the fix is usually not a faster model — it is a shorter conversation design.
Who should move, and when
| Situation | Adopt full duplex now | Stay half duplex |
|---|---|---|
| Conversational, interruption-heavy, low stakes per utterance | Yes — this is what it is for. | Only if your barge-in already works well. |
| Every action is a mutation (payments, bookings, dispatch) | Only with compensating actions built first. | Yes — the turn boundary is cheap safety. |
| Regulated call flows with scripted disclosures | Not yet — a disclosure the caller talked over is a finding. | Yes. |
| Your eval suite scores per turn | Rebuild the eval first, then adopt. | Yes, until then. |
| Long silences are structural (lookups, hold music, IVR legacy) | Model the minute cost before committing. | Possibly cheaper. |
The broader read is that OpenAI has done to the voice stack what the Agents API did to the control loop: taken a layer teams were building badly, made it excellent, and moved a decision you used to own inside a version you cannot pin. That is a good trade in both cases, and in both cases the price is your ability to attribute a regression. The defence is identical too — one frozen scenario you re-run daily, and a number you compute yourself.
FAQ
Does full duplex mean I no longer need voice activity detection?
You no longer need it to decide when the agent should reply — the model does that. You still want it for the things VAD was separately good at: detecting that a call has gone silent, measuring talk-over rates, and segmenting recordings for review.
Why does GPT-Live-1 need a separate backend model?
Because keeping a conversation flowing and reasoning carefully are different jobs with different latency budgets. Splitting them lets the voice layer stay responsive while a slower model thinks, and lets you choose how hard that model thinks per call. It also means every benchmark number is a property of the pair, not of either half.
Is $0.05 per minute cheaper than what I pay today?
It is not comparable on its own, because it buys only the voice layer. Add the backend model's tokens and any tools. The structural change is that the voice portion is now billed on elapsed time rather than on audio tokens, so silence and slow tool calls have a price they did not have before.
What breaks first when I switch an existing agent over?
Usually the tool-dispatch guard and the eval suite, in that order. Anything with the shape "wait for the user to finish, then act" has no trigger, and any scoring harness segmented by turn has no segments.
Can I keep a human-escalation path in a full-duplex call?
Yes, but the handoff point has to be declared rather than inferred. Pick an explicit condition — a confidence threshold, a caller phrase, a failed tool — and have the agent stop speaking on it, instead of waiting for a turn to end. See escalation and warm transfer.
Further reading
On this wiki:
- Full-duplex speech — the concept, and why it is not simply faster half duplex.
- Voice & realtime agents — the cascade-versus-speech-to-speech choice this sits on top of.
- Turn-taking and barge-in — the half-duplex problem in full detail.
- The latency budget — where the milliseconds actually go.
- Voice tooling and state — dialogue state as the thing commit points are built from.
- Long-lived sessions and deploys — what a call-shaped workload does to your release process.
Sources:
- Build more natural voice experiences with GPT-Live-1 in the API — OpenAI, 10 September 2026.
- OpenAI API docs — gpt-live-1 — model card, transports and session limits.
- Getting started with GPT-Live — session setup, backend models, WebRTC/WebSocket/SIP.