The feature tables for these four frameworks are nearly identical and nearly useless, because they compare the layer that is easiest to replace. What separates Pipecat, LiveKit Agents, TEN and Bolna is where each one's centre of gravity sits — a pipeline, a media server, a graph runtime, a phone line — and only the media path is expensive to change once you have callers on it. Pick on that, on telephony breadth, and on how fast the project ships; the pipeline ergonomics everyone benchmarks are being commoditised out from under all four.
At a glance
All four are free, permissively licensed in the broad sense, and capable of answering a phone. They disagree about what a voice agent fundamentally is.
| Project | Licence | Language surface | Centre of gravity |
|---|---|---|---|
| Pipecat | BSD-2-Clause | Python | A pipeline of processors; transport is a plugin. |
| LiveKit Agents | Apache-2.0 | Python, Node.js | A programmable participant inside a WebRTC room. |
| TEN Framework | Apache-2.0 with additional restrictions | C/C++, Go, Python, JS/TS | A graph of nodes, executed by a polyglot runtime. |
| Bolna | MIT | Python | A phone number with a conversation behind it. |
The axis nobody puts in the comparison table
A voice agent is four things stacked: something that carries audio between a human and your server, something that orchestrates the steps, something that bridges to the phone network, and a set of model integrations. Every comparison of these frameworks concentrates on the second and fourth — how nice the pipeline API is, how many STT vendors are wired up — because those are the parts you touch while writing the demo.
They are also the parts with the lowest switching cost. Model plugins are a day of work each and are being rewritten constantly anyway. The orchestration layer is your own code and you can port it. The media path is different: it determines your infrastructure footprint, your scaling model, your latency floor, your compliance story about where audio flows, and whether self-hosting is a checkbox or a project. Choosing it is choosing a dependency you will still have in three years.
The test that sorts these four in one question: if you wanted to move the media path in-house tomorrow, what would break? For LiveKit Agents, nothing — the media server is the open-source part. For Pipecat, you swap a transport. For TEN, you leave the path the ecosystem is built around. For Bolna, you are not moving it: the media path belongs to your telephony carrier by design.
Four projects, four answers
Pipecat — the pipeline, and the deliberate absence of a transport
Pipecat is Daily's Python framework and the most-starred of the four at around 15.5k, under BSD-2-Clause. It models an agent as a pipeline of processors through which frames flow, and it pointedly does not supply a media path: transports include Daily's WebRTC, LiveKit's WebRTC, a FastAPI WebSocket, a small built-in WebRTC transport, Vonage, and local audio. Telephony is the broadest of the four — serializers for Twilio, Vonage, Telnyx, Plivo, Exotel and Genesys — which means the same pipeline serves a web widget and an inbound number by changing the front door.
That neutrality is the whole pitch and the whole risk. You get the largest integration surface and the fastest-moving repository of the four, at roughly 12,900 commits, and you get to decide the media question yourself — which is a benefit if you have an opinion and a cost if you were hoping the framework would have one.
LiveKit Agents — the agent is a participant, not a pipeline
LiveKit Agents is Apache-2.0, around 14.2k stars, Python and Node.js, and it inverts the model: your agent is a programmable participant that joins a LiveKit room alongside human participants, running on top of LiveKit's open-source WebRTC media server. Dispatch APIs handle job scheduling and agent launch, and MCP is supported natively for tool integration.
The room abstraction is not a stylistic preference. Multi-participant scenarios — a caller, an agent, a supervisor listening in, a second agent handling a sub-task — are native rather than bolted on, and warm transfer stops being a telephony trick. Telephony is native SIP rather than a per-vendor serializer, which is the cleaner story if you are terminating your own trunks and the more opinionated one if you are not. This is the only option of the four where self-hosting the entire media path is a supported, documented configuration rather than an exercise.
TEN Framework — a graph runtime that happens to do voice
TEN, backed by Agora, is the structurally most ambitious of the four: around 11.1k stars, an agent described as a directed graph of nodes — STT, LLM, tools, memory, vision — and extensions written in C/C++, Go, Python or JS/TS rather than Python only. It ships its own components for the hard real-time parts, including a VAD and a turn-detection model aimed specifically at full-duplex dialogue.
Two things to weigh. The polyglot runtime is genuinely differentiating if you have a latency-critical component that should not be in Python, and genuinely overhead if you do not. And the licence is not plain Apache-2.0: the repository states Apache License 2.0 with additional restrictions, with the packages directory released separately under plain Apache-2.0. That is a sentence to hand to whoever approves your dependencies, before rather than after you build on it.
Bolna — telephony-first, and honest about it
Bolna is the outlier at roughly 760 stars and MIT-licensed, and the star gap misreads it. It is a Python framework built around outbound and inbound telephony — Twilio and Plivo integrated, Exotel and Vonage listed as arriving — with strong support for Indian languages and carriers, and around 2,900 commits behind it. It is not trying to be a general real-time media framework; it is trying to make a phone campaign work.
For a team whose product is outbound calling in a market the larger projects treat as secondary, that focus is worth more than a bigger plugin catalogue. For anyone else, the smaller contributor base is a real risk in a category where the model layer moves twice a year.
The subsystem they compete on is the one being deleted
Here is the awkward part for all four. The place these frameworks have invested most heavily in the last two years is turn detection: knowing when the human has finished speaking so the agent may reply. TEN ships a turn-detection model as a headline component. LiveKit and Pipecat both ship semantic end-of-turn detection. It is the hardest problem in the category and the one with the most visible quality delta.
It is also the problem a full-duplex model makes moot. OpenAI's GPT-Live-1, in the API since 10 September 2026, listens while it speaks and decides for itself when to talk — there is no end-of-turn event for a framework to compute, and no threshold for it to tune. Every framework in this comparison will keep a full-duplex adapter; the question is what is left of their differentiation once the adapter is the common path.
What survives is exactly the list this post argues for: who owns the media path, how many carriers you can reach, how observable the session is, and how quickly the project ships when the model layer changes again. What does not survive is pipeline elegance. If you are choosing today on which API felt nicest in a weekend prototype, you are optimising the part with the shortest half-life.
Vocode is the cautionary case and it is recent. It was a credible member of this list, MIT-licensed with first-class Twilio and Vonage support, and it stalled — minimal commits for well over a year and an architecture that predates both speech-to-speech models and sub-500 ms pipelines. Nothing about it was wrong when it was chosen. In this category, a framework that stops shipping does not degrade gracefully; it becomes a rewrite the first time the model layer moves.
Licence and maintenance are buying axes here, not footnotes
Three of the four are licences your legal team will approve without reading: BSD-2-Clause for Pipecat, Apache-2.0 for LiveKit Agents, MIT for Bolna. TEN's "Apache-2.0 with additional restrictions" is the one that needs an actual read, and the honest framing is not that it is disqualifying — it is that the cost of finding out is five minutes now and a migration later.
Maintenance velocity deserves the same weight. The useful signal is not stars; the top three here sit within about 30% of each other, and Pipecat's lead partly reflects how the organisation splits its repositories. The signal is commit cadence against a moving model layer. Two full-duplex-capable voice models and a wave of new STT and TTS providers landed in the last twelve months, and a framework's value is almost entirely in having already absorbed them.
When to pick which
| Situation | Pipecat | LiveKit Agents | TEN / Bolna |
|---|---|---|---|
| Phone-first, many carriers | Yes — six serializers. | Yes, if you terminate SIP yourself. | Bolna, if the market is India. |
| Browser or app, multi-participant | Workable. | Yes — the room is the model. | TEN, if Agora is already there. |
| Must self-host the whole media path | Only with a transport you host. | Yes — the media server is open source. | No. |
| A latency-critical component outside Python | No. | No. | TEN — this is its reason to exist. |
| Widest provider choice, fastest to prototype | Yes. | Close second. | No. |
| Dependency review is strict | BSD-2-Clause. | Apache-2.0. | Read TEN's licence first; Bolna is MIT. |
The default recommendation, stated plainly: if you do not already own a media path, take LiveKit Agents, because the decision you are least able to revisit is the one it makes for you and makes well. If you do own one — a Daily account, a Twilio estate, an existing WebRTC deployment — take Pipecat, because its refusal to have an opinion about transport is exactly what you want. Reach for TEN when the polyglot runtime answers a question you actually have, and for Bolna when telephony in its home market is the product rather than the plumbing.
Whichever you pick, the piece to build yourself is the same one all four leave thin: evaluation. A framework will get you a call that connects. Nothing in any of these repositories tells you whether the call went well.
FAQ
Is Pipecat's higher star count a reason to choose it?
Not on its own. The top three are within roughly 30% of each other, which is inside the noise created by how each organisation splits its repositories. Commit cadence and how fast a new model provider gets a plugin are far better signals.
Can I use Pipecat and LiveKit together?
Yes, and it is a common shape. LiveKit is one of Pipecat's supported transports, so you can run a Pipecat pipeline over a LiveKit media path and get the pipeline ergonomics with the room model underneath.
Does full duplex make these frameworks unnecessary?
No. It removes one subsystem — turn detection — and leaves transport, telephony bridging, session state, tool dispatch, observability and provider plumbing untouched. It changes what you should compare them on, not whether you need one.
What exactly is the issue with TEN's licence?
The repository is Apache License 2.0 with additional restrictions rather than plain Apache-2.0, with the packages directory released separately under unmodified Apache-2.0. Whether that matters depends on your distribution model, which is why it should be read before adoption rather than during a review.
Where does Vapi, Retell or ElevenLabs fit against these?
They are the managed tier, not alternatives at this layer — you rent the whole stack rather than assembling one. That comparison is a separate post; the trade is the usual one between speed to launch and control over the media path.
Further reading
On this wiki:
- Realtime architecture — the reference shape production voice stacks converge on.
- Telephony and PSTN integration — what changes when the caller is on a phone.
- Turn-taking and barge-in — the subsystem this post says is being commoditised.
- Full-duplex speech — why the turn boundary is disappearing.
- Evaluating voice agents — the part none of the four ships.
- Agent frameworks — how to judge this class of dependency in general.
Project sources:
- pipecat-ai/pipecat — transports, telephony serializers, BSD-2-Clause.
- livekit/agents — programmable participants, dispatch APIs, Apache-2.0.
- TEN-framework/ten-framework — graph runtime, TEN VAD and turn detection, licence terms.
- bolna-ai/bolna — telephony-first conversational agents, MIT.