AI Blog

Pipecat vs LiveKit Agents vs TEN vs Bolna: buy the media path, not the pipeline

Four open-source voice frameworks that look interchangeable on a feature table have their centres of gravity in four different columns — the runtime, the media server, the graph, the phone line — and only one of those is expensive to change later. The pipeline ergonomics everyone benchmarks are also the part a full-duplex model is busy commoditising, so pick on transport ownership, telephony breadth and maintenance velocity, and read TEN’s licence before you ship.

By Agentic AI Wiki 13 min read

The feature tables for these four frameworks are nearly identical and nearly useless, because they compare the layer that is easiest to replace. What separates Pipecat, LiveKit Agents, TEN and Bolna is where each one's centre of gravity sits — a pipeline, a media server, a graph runtime, a phone line — and only the media path is expensive to change once you have callers on it. Pick on that, on telephony breadth, and on how fast the project ships; the pipeline ergonomics everyone benchmarks are being commoditised out from under all four.

At a glance

All four are free, permissively licensed in the broad sense, and capable of answering a phone. They disagree about what a voice agent fundamentally is.

ProjectLicenceLanguage surfaceCentre of gravity
Pipecat BSD-2-Clause Python A pipeline of processors; transport is a plugin.
LiveKit Agents Apache-2.0 Python, Node.js A programmable participant inside a WebRTC room.
TEN Framework Apache-2.0 with additional restrictions C/C++, Go, Python, JS/TS A graph of nodes, executed by a polyglot runtime.
Bolna MIT Python A phone number with a conversation behind it.
GitHub stars by project, September 2026 Horizontal bar chart of four repositories. Pipecat leads at roughly 15,500 stars, LiveKit Agents follows at roughly 14,200, TEN Framework at roughly 11,100, and Bolna trails far behind at roughly 760. The top three bars are close together; the fourth is an order of magnitude shorter. GitHub stars (thousands) Pipecat 15.5k LiveKit Agents 14.2k TEN Framework 11.1k Bolna 0.76k 4k 8k 12k 16k September 2026, rounded. Star counts reflect repository packaging as much as adoption.
The top three are within one repackaging of each other. The gap that matters is the fourth bar.
Where each voice framework leans hardest Feature matrix with four project rows and five axis columns: media path ownership, telephony breadth, polyglot extensions, maintenance velocity and licence simplicity. Pipecat is weak on media path, strong on telephony, weak on polyglot, strong on maintenance and strong on licence. LiveKit Agents is strong on media path, strong on telephony, medium on polyglot, strong on maintenance and strong on licence. TEN Framework is medium on media path, weak on telephony, strong on polyglot, medium on maintenance and weak on licence simplicity. Bolna is weak on media path, medium on telephony, weak on polyglot, medium on maintenance and strong on licence. Where each project leans hardest Media path Telephony Polyglot Maintenance Licence Pipecat Weak bring a transport Strong 6 serializers Weak Python only Strong ~12.9k commits Strong BSD-2-Clause LiveKit Agents Strong own WebRTC SFU Strong native SIP Medium Python + Node Strong active Strong Apache-2.0 TEN Framework Medium Agora-shaped Weak thin coverage Strong C/Go/Py/JS Medium steady Weak added restrictions Bolna Weak carrier owns it Medium Twilio, Plivo Weak Python only Medium ~2.9k commits Strong MIT Strong lean Medium Weak
Each row has a different shape. None of them is better on every axis, which is the useful finding.

The axis nobody puts in the comparison table

Each framework's centre of gravity, by layer Four rows, one per project, across four columns: transport and media, agent runtime, telephony, and model plugins. Pipecat's solid accent cell sits in the agent runtime column, with transport left to you. LiveKit Agents' solid accent cell sits in the transport and media column, on its own self-hostable WebRTC media server. TEN Framework's solid accent cell sits in the agent runtime column as a polyglot node graph. Bolna's solid accent cell sits in the telephony column. Lighter cells are supplied but swappable; neutral cells are ones you bring yourself. Which layer each project is actually selling Transport / media Agent runtime Telephony Model plugins Pipecat Bring your own Daily, LiveKit, WS Processor pipeline frames flow through Six serializers Twilio → Genesys Broadest catalogue swap per stage LiveKit Agents Own WebRTC SFU self-hostable Room participant dispatch APIs Native SIP your own trunks Plugins + MCP tools built in TEN Framework Agora, or bring one ecosystem default Node graph C / Go / Python / JS Thin not the focus Extensions own VAD and turn model Bolna Carrier's not yours to move Python agent campaign-shaped Twilio, Plivo outbound-first Focused set regional providers centre of gravity supplied, swappable you bring it
Four projects, four different columns. The one in the left column is the one you cannot casually replace later.

A voice agent is four things stacked: something that carries audio between a human and your server, something that orchestrates the steps, something that bridges to the phone network, and a set of model integrations. Every comparison of these frameworks concentrates on the second and fourth — how nice the pipeline API is, how many STT vendors are wired up — because those are the parts you touch while writing the demo.

They are also the parts with the lowest switching cost. Model plugins are a day of work each and are being rewritten constantly anyway. The orchestration layer is your own code and you can port it. The media path is different: it determines your infrastructure footprint, your scaling model, your latency floor, your compliance story about where audio flows, and whether self-hosting is a checkbox or a project. Choosing it is choosing a dependency you will still have in three years.

The test that sorts these four in one question: if you wanted to move the media path in-house tomorrow, what would break? For LiveKit Agents, nothing — the media server is the open-source part. For Pipecat, you swap a transport. For TEN, you leave the path the ecosystem is built around. For Bolna, you are not moving it: the media path belongs to your telephony carrier by design.

Four projects, four answers

Pipecat — the pipeline, and the deliberate absence of a transport

Pipecat is Daily's Python framework and the most-starred of the four at around 15.5k, under BSD-2-Clause. It models an agent as a pipeline of processors through which frames flow, and it pointedly does not supply a media path: transports include Daily's WebRTC, LiveKit's WebRTC, a FastAPI WebSocket, a small built-in WebRTC transport, Vonage, and local audio. Telephony is the broadest of the four — serializers for Twilio, Vonage, Telnyx, Plivo, Exotel and Genesys — which means the same pipeline serves a web widget and an inbound number by changing the front door.

That neutrality is the whole pitch and the whole risk. You get the largest integration surface and the fastest-moving repository of the four, at roughly 12,900 commits, and you get to decide the media question yourself — which is a benefit if you have an opinion and a cost if you were hoping the framework would have one.

LiveKit Agents — the agent is a participant, not a pipeline

LiveKit Agents is Apache-2.0, around 14.2k stars, Python and Node.js, and it inverts the model: your agent is a programmable participant that joins a LiveKit room alongside human participants, running on top of LiveKit's open-source WebRTC media server. Dispatch APIs handle job scheduling and agent launch, and MCP is supported natively for tool integration.

The room abstraction is not a stylistic preference. Multi-participant scenarios — a caller, an agent, a supervisor listening in, a second agent handling a sub-task — are native rather than bolted on, and warm transfer stops being a telephony trick. Telephony is native SIP rather than a per-vendor serializer, which is the cleaner story if you are terminating your own trunks and the more opinionated one if you are not. This is the only option of the four where self-hosting the entire media path is a supported, documented configuration rather than an exercise.

TEN Framework — a graph runtime that happens to do voice

TEN, backed by Agora, is the structurally most ambitious of the four: around 11.1k stars, an agent described as a directed graph of nodes — STT, LLM, tools, memory, vision — and extensions written in C/C++, Go, Python or JS/TS rather than Python only. It ships its own components for the hard real-time parts, including a VAD and a turn-detection model aimed specifically at full-duplex dialogue.

Two things to weigh. The polyglot runtime is genuinely differentiating if you have a latency-critical component that should not be in Python, and genuinely overhead if you do not. And the licence is not plain Apache-2.0: the repository states Apache License 2.0 with additional restrictions, with the packages directory released separately under plain Apache-2.0. That is a sentence to hand to whoever approves your dependencies, before rather than after you build on it.

Bolna — telephony-first, and honest about it

Bolna is the outlier at roughly 760 stars and MIT-licensed, and the star gap misreads it. It is a Python framework built around outbound and inbound telephony — Twilio and Plivo integrated, Exotel and Vonage listed as arriving — with strong support for Indian languages and carriers, and around 2,900 commits behind it. It is not trying to be a general real-time media framework; it is trying to make a phone campaign work.

For a team whose product is outbound calling in a market the larger projects treat as secondary, that focus is worth more than a bigger plugin catalogue. For anyone else, the smaller contributor base is a real risk in a category where the model layer moves twice a year.

The subsystem they compete on is the one being deleted

Here is the awkward part for all four. The place these frameworks have invested most heavily in the last two years is turn detection: knowing when the human has finished speaking so the agent may reply. TEN ships a turn-detection model as a headline component. LiveKit and Pipecat both ship semantic end-of-turn detection. It is the hardest problem in the category and the one with the most visible quality delta.

It is also the problem a full-duplex model makes moot. OpenAI's GPT-Live-1, in the API since 10 September 2026, listens while it speaks and decides for itself when to talk — there is no end-of-turn event for a framework to compute, and no threshold for it to tune. Every framework in this comparison will keep a full-duplex adapter; the question is what is left of their differentiation once the adapter is the common path.

What survives is exactly the list this post argues for: who owns the media path, how many carriers you can reach, how observable the session is, and how quickly the project ships when the model layer changes again. What does not survive is pipeline elegance. If you are choosing today on which API felt nicest in a weekend prototype, you are optimising the part with the shortest half-life.

Vocode is the cautionary case and it is recent. It was a credible member of this list, MIT-licensed with first-class Twilio and Vonage support, and it stalled — minimal commits for well over a year and an architecture that predates both speech-to-speech models and sub-500 ms pipelines. Nothing about it was wrong when it was chosen. In this category, a framework that stops shipping does not degrade gracefully; it becomes a rewrite the first time the model layer moves.

Licence and maintenance are buying axes here, not footnotes

Three of the four are licences your legal team will approve without reading: BSD-2-Clause for Pipecat, Apache-2.0 for LiveKit Agents, MIT for Bolna. TEN's "Apache-2.0 with additional restrictions" is the one that needs an actual read, and the honest framing is not that it is disqualifying — it is that the cost of finding out is five minutes now and a migration later.

Maintenance velocity deserves the same weight. The useful signal is not stars; the top three here sit within about 30% of each other, and Pipecat's lead partly reflects how the organisation splits its repositories. The signal is commit cadence against a moving model layer. Two full-duplex-capable voice models and a wave of new STT and TTS providers landed in the last twelve months, and a framework's value is almost entirely in having already absorbed them.

When to pick which

SituationPipecatLiveKit AgentsTEN / Bolna
Phone-first, many carriers Yes — six serializers. Yes, if you terminate SIP yourself. Bolna, if the market is India.
Browser or app, multi-participant Workable. Yes — the room is the model. TEN, if Agora is already there.
Must self-host the whole media path Only with a transport you host. Yes — the media server is open source. No.
A latency-critical component outside Python No. No. TEN — this is its reason to exist.
Widest provider choice, fastest to prototype Yes. Close second. No.
Dependency review is strict BSD-2-Clause. Apache-2.0. Read TEN's licence first; Bolna is MIT.

The default recommendation, stated plainly: if you do not already own a media path, take LiveKit Agents, because the decision you are least able to revisit is the one it makes for you and makes well. If you do own one — a Daily account, a Twilio estate, an existing WebRTC deployment — take Pipecat, because its refusal to have an opinion about transport is exactly what you want. Reach for TEN when the polyglot runtime answers a question you actually have, and for Bolna when telephony in its home market is the product rather than the plumbing.

Whichever you pick, the piece to build yourself is the same one all four leave thin: evaluation. A framework will get you a call that connects. Nothing in any of these repositories tells you whether the call went well.

FAQ

Is Pipecat's higher star count a reason to choose it?

Not on its own. The top three are within roughly 30% of each other, which is inside the noise created by how each organisation splits its repositories. Commit cadence and how fast a new model provider gets a plugin are far better signals.

Can I use Pipecat and LiveKit together?

Yes, and it is a common shape. LiveKit is one of Pipecat's supported transports, so you can run a Pipecat pipeline over a LiveKit media path and get the pipeline ergonomics with the room model underneath.

Does full duplex make these frameworks unnecessary?

No. It removes one subsystem — turn detection — and leaves transport, telephony bridging, session state, tool dispatch, observability and provider plumbing untouched. It changes what you should compare them on, not whether you need one.

What exactly is the issue with TEN's licence?

The repository is Apache License 2.0 with additional restrictions rather than plain Apache-2.0, with the packages directory released separately under unmodified Apache-2.0. Whether that matters depends on your distribution model, which is why it should be read before adoption rather than during a review.

Where does Vapi, Retell or ElevenLabs fit against these?

They are the managed tier, not alternatives at this layer — you rent the whole stack rather than assembling one. That comparison is a separate post; the trade is the usual one between speed to launch and control over the media path.

Further reading

On this wiki:

Project sources: