AI Blog

Tagged: realtime

← Back to AI Blog

10 min read

Pipecat vs LiveKit Agents vs TEN vs Bolna: buy the media path, not the pipeline

Four open-source voice frameworks that look interchangeable on a feature table have their centres of gravity in four different columns — the runtime, the media server, the graph, the phone line — and only one of those is expensive to change later. The pipeline ergonomics everyone benchmarks are also the part a full-duplex model is busy commoditising, so pick on transport ownership, telephony breadth and maintenance velocity, and read TEN’s licence before you ship.

10 min read

Full duplex deletes the turn — and the turn was your commit point

GPT-Live-1 landed in the API on 10 September and listens while it speaks, which reads as a naturalness upgrade and is actually a schema change. End-of-turn was the event your voice agent used to decide when to call a tool, when to write a log line, when to run a guardrail and when to stop the meter — and a full-duplex model never fires it. The fix is not a better threshold; it is naming your own commit points and pricing a meter that now runs on wall clock instead of speech.

9 min read

Coval vs Hamming vs Cekura vs Bluejay: You Are Buying a Simulated Caller

Four platforms will run thousands of test calls against your voice agent, and the number they advertise — concurrency — is the axis that matters least. What separates them is where the caller on the other end comes from, because that sets the ceiling on what any of these evals can tell you.

9 min read

ElevenLabs vs Cartesia vs Deepgram vs Rime: buy the tail, not the average

These four advertise time-to-first-audio between 40 and 200 ms, and an independent harness measures their cloud medians at 188 to 313 ms — but the number that breaks a phone call is the spread, not the median, and one vendor's jitter is nearly four times another's. Price moves about 2.5× across the field and predicts neither. The tail is bought with deployment.

8 min read

Deepgram vs AssemblyAI vs ElevenLabs vs Speechmatics: You Are Buying a Turn Detector

Word error rate is close to settled between the four, and a couple of points of it lands on words your intent classifier ignores. The slice that decides whether a voice agent feels human is end-of-turn detection — several times larger than the transcription latency beneath it, and the one thing the four providers genuinely disagree about.

13 min read

ElevenLabs vs Vapi vs Retell vs OpenAI gpt-realtime: Four Bets on How Your Agent Should Talk Back

Voice is now the interface most agents will spend the most time in — and four platforms have made architecturally opposite bets on how to wire speech, language, and tool-use into one round-trip. The right pick depends less on TTS voice quality than on whether you control the audio path, the model, or just the prompt.