All four of these platforms ingest OpenTelemetry, which is why "it supports OTel" has stopped being a useful answer: the vocabulary that would actually make a trace portable — the gen_ai.* attributes — is still, in late 2026, entirely at Development stability and now lives in its own repository, so there is no frozen contract for anyone to conform to. That matters more here than in ordinary observability because production traces are the one asset in your agent stack you cannot re-create. You can re-run an eval, rewrite a prompt, retrain a judge. You cannot re-observe last quarter's traffic. So choose on who owns the write path and the bulk read path, and treat the licence badge as the orthogonal fact it is.
At a glance
Four platforms that look interchangeable on a feature grid and are not interchangeable at all on exit.
| Platform | Licence | Deployment shape | Strongest at |
|---|---|---|---|
| Langfuse | MIT core; commercial licence for governance features | Cloud, or self-host on every tier | Owning both the code and the store without a negotiation. |
| Phoenix (Arize) | Elastic License 2.0 — source-available, not OSI | One process: pip, Docker or Helm | Getting a retrieval-shaped trace readable in an afternoon. |
| Braintrust | Proprietary | Cloud; hybrid data plane in your account on enterprise | Evaluation as the primary workflow rather than a tab. |
| LangSmith | Proprietary | Cloud; self-host on the enterprise plan | Teams already inside the LangChain and LangGraph ergonomics. |
Read the licence column and the deployment column as separate facts, because they come apart in both directions. Phoenix's licence is not OSI-approved and it still hands you the database. Braintrust is proprietary top to bottom and will still put the span bytes in your own object storage. That combination is unusual enough in infrastructure that most teams mis-rank the four on instinct.
"Supports OpenTelemetry" names the protocol, not the contract
OTLP is a wire format. It says how a span is serialised and shipped; it says nothing about what the fields are called. Portability of a trace lives entirely in the second question, and in 2026 the answer is unsettled in a way that vendor comparison tables do not show. Every gen_ai.* attribute, span, metric and event in the OpenTelemetry registry carries the Development stability badge. In June 2026 the generative-AI conventions were separated out of open-telemetry/semantic-conventions into a dedicated semantic-conventions-genai repository, which now also carries MCP and provider-specific conventions.
Development stability is a real status with real consequences: an attribute at that level can be renamed or removed between releases without the deprecation window a stable convention would get. So when four vendors all say they support the OTel GenAI conventions, they are each conforming to a moving target, at different commits, with different amounts of their own vocabulary mixed in. Phoenix does not even pretend otherwise — it is native to OpenInference, Arize's own set of OTel conventions with explicit span kinds for chain, agent, retriever, embedding, tool, LLM and reranker. That is arguably the more honest position, and it is also a vocabulary lock-in with a different name.
The practical test, which takes twenty minutes and beats any comparison table: send one agent run with three tool calls and a judge score into a trial account. Then try to get it back out. Not "is there an export button" — can you retrieve the attributes you sent, under the names you sent them, in bulk, without pagination limits that make a year of history a multi-week job? Everything else in this post is downstream of that answer. See OTel GenAI semantic conventions for what the vocabulary does and does not yet guarantee.
Where the bytes land
Langfuse — the only one where code and store are both yours by default
Langfuse open-sourced its product features under MIT in June 2025, and self-hosting is available on every tier with no usage fee. Tracing, prompt management, datasets, LLM-as-a-judge evaluation, the playground and annotation queues ship in the free distribution. What is commercially licensed is a short, specific list of governance features — project-level role-based access control, SCIM, audit logs, protected prompt labels, retention policies — behind a paid licence key. The spans land in Postgres and ClickHouse that you run, which makes export a query rather than an API budget.
The structural consequence is that leaving Langfuse is a schema-translation problem, not a data-retrieval problem. That is the cheapest shape of exit available in this category.
Phoenix — source-available, single process, OpenInference-native
Phoenix runs as one process via pip, Docker or Helm, with no per-event caps when self-hosted, which makes it the fastest of the four to get a trace rendered on a laptop. Its licence is Elastic License 2.0: broad internal use and modification are permitted, offering Phoenix as a hosted managed service is not, and it is not an OSI-approved open-source licence. If your compliance review has a binary "is it open source" field, that field will be answered wrongly in both directions here — the licence restricts a thing you were not going to do, and grants the thing you actually care about.
Braintrust — proprietary engine, your object storage
Braintrust is eval-first: the tracing is in service of regression testing rather than the other way round, and that shows in where its ergonomics are good. Its hybrid deployment, available on enterprise plans, uses Terraform to put the data plane — API, Postgres, Redis, object storage and Brainstore — in your own cloud account while Braintrust hosts the control plane. Brainstore is its own Rust engine, writing spans to object storage and indexing them with the open-source Tantivy library.
So the bytes can be in your account while the thing that can read them efficiently is not yours. That is a genuinely different risk profile from both the open-core and the pure-SaaS cases: your data-residency and sovereignty questions get good answers, and your exit question gets a worse one than the storage location suggests.
LangSmith — managed by default, self-hosted by contract
LangSmith is proprietary, usage-priced, and self-hostable only on the enterprise plan. For a team already writing LangChain and LangGraph, the integration cost is close to zero and the trace views match the abstractions in the code, which is a real and underrated advantage — a trace that mirrors your own control flow gets read, and one that does not, does not. The trade is that on every plan below enterprise, both the code and the store are someone else's, so your historical traces exist at a vendor's pagination limits.
The licence tells you about the code, not about the data
It is worth stating the failure mode directly, because it is the most common mistake in this category. Teams rank these four on how open the licence is, conclude that the MIT option is the portable one and the proprietary ones are traps, and then make a choice that does not survive contact with the actual migration.
Sort by exit cost instead and the ordering changes. Phoenix, under a licence that is not open source by the OSI definition, gives you a database you can dump. Braintrust, proprietary throughout, can put the span bytes in your S3 bucket. LangSmith's enterprise self-host puts the whole thing in your VPC under a commercial agreement. Meanwhile an MIT licence on a hosted plan you never self-host buys you the theoretical right to run code you are not running — which is worth something as insurance against the vendor disappearing, and nothing at all for getting last quarter's traces into a different tool next month.
# Rank on these, in this order 1 can I run a bulk read of historical spans, un-paginated? 2 do the stored attribute names come from my code or the vendor's SDK? 3 where do the bytes physically sit, and under whose contract? 4 what is the licence? # Note what is NOT on the list dashboard quality — you will look at it twice a week number of integrations — you need two of them judge library size — you will write your own judges
None of which makes the licence worthless: a commercial licence that can be withdrawn is a vendor-risk entry, and the right place for it is third-party and vendor risk, not the portability column.
Own the shim and the choice gets smaller
The most useful thing in this post is not a ranking. It is that the portability decision happens upstream of the platform, in roughly two hundred lines of code you write once: a thin layer between your agent loop and whichever SDK you export with, where you decide the attribute names for the things you will still care about in a year — run id, task class, tool name, retry index, token counts split by cache state, judge verdict and version, human override. Emit those under names you chose, map them to the vendor's expectations at the boundary, and the platform becomes a renderer over your schema rather than the owner of it.
Two properties follow, and they are what make the shim worth the afternoon. Dual-writing during a migration becomes a configuration change rather than a project, so you can run two platforms side by side for a month and compare them on your own traffic instead of on a demo dataset. And the churn in the gen_ai.* conventions becomes someone else's problem: when an attribute is renamed upstream, you change one mapping table, and every dashboard and saved query you own keeps working. Pin the convention version you build against and move it deliberately, the same discipline as anything else your vendor can change without a deploy.
Two things the shim does not fix. It does not retroactively port history you already accumulated in a vendor's model, which is the argument for doing this before the first production run rather than after the first renewal. And it does not reduce volume: whatever you choose, trace cost scales with traffic and the lever is sampling, so decide retention and sampling policy at the same time — see trace sampling and retention, and keep redaction in the shim where it belongs rather than in a vendor setting.
When to pick which
| Situation | Pick | Because |
|---|---|---|
| Regulated data, no vendor contract in sight | Langfuse | MIT core self-hosts on every tier; the store is yours without a negotiation. |
| You need a trace readable today | Phoenix | One process, no event caps, retrieval-shaped span kinds out of the box. |
| Your bottleneck is regression testing, not debugging | Braintrust | Eval is the primary object; the hybrid data plane answers residency. |
| LangChain or LangGraph throughout, small team | LangSmith | Zero integration cost and traces that mirror your own control flow. |
| You expect to change your mind within a year | Any, plus the shim | The shim is what makes the first answer reversible. |
And the honest summary of the field: none of these four is bad, the dashboards are converging, and the difference that will matter to you in eighteen months is whether the traces you are accumulating right now are written in a vocabulary you control. That is a decision you are making today whether or not you are making it deliberately.
FAQ
If all four accept OTLP, can I just switch later?
You can switch where new spans go in an afternoon. What does not switch is the history: each platform stores spans in its own model, so moving a year of traces means a bulk read through whatever export path exists, plus a schema translation. That is why the question to ask in a trial is about bulk read, not about ingest.
Are the OpenTelemetry GenAI conventions safe to standardise on?
Safe to adopt, not safe to treat as frozen. Every gen_ai.* attribute still carries Development stability, and the conventions now live in a dedicated repository. Instrument against them, pin the version you build on, and keep a mapping layer you own so a rename upstream is a one-file change.
Is Phoenix open source?
Source-available. Elastic License 2.0 permits broad internal use and modification but not offering Phoenix as a hosted managed service, and it is not OSI-approved. For most teams that distinction changes nothing operationally; for a procurement checklist with a binary field, it changes the answer.
Does Braintrust's hybrid deployment mean I own my data?
You own the bytes and their location — the data plane, including Brainstore, runs in your cloud account via Terraform on enterprise plans. The engine that reads them efficiently remains proprietary, so residency is answered well and exit is answered less well than the storage location implies.
What is the one thing to build before choosing?
The instrumentation shim: your own attribute names for run id, task class, tool name, retry index, token counts by cache state, judge verdict and version. Roughly two hundred lines, written once, and it converts the platform choice from a one-way door into a configuration value.
Further reading
On this wiki:
- OTel GenAI Semantic Conventions — what the vocabulary covers and what it still does not promise.
- Tracing & Observability for Agents — what a useful agent span contains in the first place.
- Trace Sampling & Retention — the cost lever none of these four pulls for you.
- Eval-Driven Agent Development — the workflow Braintrust is organised around.
- Agent Observability — the concept page, if you are arriving without the vocabulary.
Project sources:
- Langfuse — MIT core, self-hosting on every tier, and the enterprise governance list.
- Arize Phoenix — Elastic License 2.0, OpenInference span kinds.
- Braintrust — OTLP ingest, hybrid data plane, Brainstore.
- LangSmith — plans, self-hosting availability.
- open-telemetry/semantic-conventions-genai — the dedicated GenAI conventions repository.