AI Blog

LangSmith vs Langfuse vs Braintrust vs Phoenix

All four ingest OpenTelemetry, so "OTel support" decides nothing — the vocabulary that would make a trace portable is still entirely at Development stability. Pick on who owns the write path and the bulk read path, because production traces are the one asset you cannot re-create, and the licence badge is orthogonal to whether you can get them back.

By Agentic AI Wiki 15 min read

All four of these platforms ingest OpenTelemetry, which is why "it supports OTel" has stopped being a useful answer: the vocabulary that would actually make a trace portable — the gen_ai.* attributes — is still, in late 2026, entirely at Development stability and now lives in its own repository, so there is no frozen contract for anyone to conform to. That matters more here than in ordinary observability because production traces are the one asset in your agent stack you cannot re-create. You can re-run an eval, rewrite a prompt, retrain a judge. You cannot re-observe last quarter's traffic. So choose on who owns the write path and the bulk read path, and treat the licence badge as the orthogonal fact it is.

At a glance

Four platforms that look interchangeable on a feature grid and are not interchangeable at all on exit.

PlatformLicenceDeployment shapeStrongest at
LangfuseMIT core; commercial licence for governance featuresCloud, or self-host on every tierOwning both the code and the store without a negotiation.
Phoenix (Arize)Elastic License 2.0 — source-available, not OSIOne process: pip, Docker or HelmGetting a retrieval-shaped trace readable in an afternoon.
BraintrustProprietaryCloud; hybrid data plane in your account on enterpriseEvaluation as the primary workflow rather than a tab.
LangSmithProprietaryCloud; self-host on the enterprise planTeams already inside the LangChain and LangGraph ergonomics.

Read the licence column and the deployment column as separate facts, because they come apart in both directions. Phoenix's licence is not OSI-approved and it still hands you the database. Braintrust is proprietary top to bottom and will still put the span bytes in your own object storage. That combination is unusual enough in infrastructure that most teams mis-rank the four on instinct.

"Supports OpenTelemetry" names the protocol, not the contract

Where a span's vocabulary is decided, and where the bytes land A span starts in your application code, where an instrumentation library names its attributes. That naming step is the only place portability is decided, and it is code you can own. Below it, the OTLP wire protocol carries the span to a platform. All four platforms accept OTLP, so the protocol is not the differentiator. Each platform then writes the span into its own store under its own vocabulary: Langfuse into Postgres and ClickHouse you can run, Phoenix into a self-hosted database under the OpenInference conventions, Braintrust into Brainstore which can sit in your own cloud account, LangSmith into a managed store unless you are on an enterprise self-hosted plan. The bulk read path back out is the axis that decides whether last quarter's traces are still yours. The protocol is shared. The vocabulary and the store are not. Your agent code one run · tool calls · retries · judge scores the only thing that knows what mattered Instrumentation — where attribute names are chosen gen_ai.* · openinference.* · vendor SDK fields THIS is the portability decision, and it is yours to own OTLP over HTTP or gRPC all four accept it — so it settles nothing BELOW THIS LINE, EACH PLATFORM REWRITES THE SPAN INTO ITS OWN STORE Langfuse MIT core, self-host on every tier Postgres + ClickHouse you operate CODE YOURS · DATA YOURS Phoenix Elastic 2.0, source- available, pip/Docker/Helm OpenInference span kinds, its own database DATA YOURS · NO RESALE Braintrust proprietary; hybrid data plane on enterprise Brainstore on object storage in your account BYTES YOURS · ENGINE NOT LangSmith proprietary SaaS; self-host on enterprise managed store on every other plan NEITHER, BY DEFAULT You can re-run an eval and rewrite a prompt. You cannot re-observe last quarter's production traffic. So read the bottom row as a question about exit, not about ideology.
The transport is shared. The vocabulary and the store are where the four actually differ.

OTLP is a wire format. It says how a span is serialised and shipped; it says nothing about what the fields are called. Portability of a trace lives entirely in the second question, and in 2026 the answer is unsettled in a way that vendor comparison tables do not show. Every gen_ai.* attribute, span, metric and event in the OpenTelemetry registry carries the Development stability badge. In June 2026 the generative-AI conventions were separated out of open-telemetry/semantic-conventions into a dedicated semantic-conventions-genai repository, which now also carries MCP and provider-specific conventions.

Development stability is a real status with real consequences: an attribute at that level can be renamed or removed between releases without the deprecation window a stable convention would get. So when four vendors all say they support the OTel GenAI conventions, they are each conforming to a moving target, at different commits, with different amounts of their own vocabulary mixed in. Phoenix does not even pretend otherwise — it is native to OpenInference, Arize's own set of OTel conventions with explicit span kinds for chain, agent, retriever, embedding, tool, LLM and reranker. That is arguably the more honest position, and it is also a vocabulary lock-in with a different name.

The practical test, which takes twenty minutes and beats any comparison table: send one agent run with three tool calls and a judge score into a trial account. Then try to get it back out. Not "is there an export button" — can you retrieve the attributes you sent, under the names you sent them, in bulk, without pagination limits that make a year of history a multi-week job? Everything else in this post is downstream of that answer. See OTel GenAI semantic conventions for what the vocabulary does and does not yet guarantee.

Where the bytes land

Four platforms against the four axes that are actually contested A matrix with four platforms as rows and four axes as columns. On self-hosting, Langfuse is strong because the MIT core self-hosts on every tier and Phoenix is strong as a single source-available process, while Braintrust offers a hybrid data plane only on enterprise plans and LangSmith self-hosts only on enterprise. On where the bytes sit, Langfuse and Phoenix put them in databases you operate, Braintrust can put them in your own object storage behind a proprietary engine, and LangSmith keeps them managed by default. On span vocabulary, Langfuse and LangSmith ingest the OpenTelemetry generative-AI attributes, Phoenix is native to the OpenInference conventions, and Braintrust maps OTLP into its own model. On eval depth, Braintrust is strongest, Langfuse and Phoenix are capable, and LangSmith is tied to its own framework ergonomics. Where each one leans hardest SELF-HOST WHERE BYTES SIT SPAN VOCABULARY EVAL DEPTH Langfuse every tier, MIT your Postgres gen_ai.* ingest judges + datasets Phoenix one process your database OpenInference RAG-shaped Braintrust hybrid, enterprise your object store own model eval-first LangSmith enterprise only managed gen_ai.* ingest framework-native strong capable weakest of the four No row is bad. Column three is the one every vendor answers with the same word and means four things by. Ingesting OTLP is not the same as storing, querying and re-emitting the gen_ai.* attributes you sent. Column two is the only one with a migration cost attached, which is why it should move first in a decision.
Column two is the only one with a migration cost attached to it.

Langfuse — the only one where code and store are both yours by default

Langfuse open-sourced its product features under MIT in June 2025, and self-hosting is available on every tier with no usage fee. Tracing, prompt management, datasets, LLM-as-a-judge evaluation, the playground and annotation queues ship in the free distribution. What is commercially licensed is a short, specific list of governance features — project-level role-based access control, SCIM, audit logs, protected prompt labels, retention policies — behind a paid licence key. The spans land in Postgres and ClickHouse that you run, which makes export a query rather than an API budget.

The structural consequence is that leaving Langfuse is a schema-translation problem, not a data-retrieval problem. That is the cheapest shape of exit available in this category.

Phoenix — source-available, single process, OpenInference-native

Phoenix runs as one process via pip, Docker or Helm, with no per-event caps when self-hosted, which makes it the fastest of the four to get a trace rendered on a laptop. Its licence is Elastic License 2.0: broad internal use and modification are permitted, offering Phoenix as a hosted managed service is not, and it is not an OSI-approved open-source licence. If your compliance review has a binary "is it open source" field, that field will be answered wrongly in both directions here — the licence restricts a thing you were not going to do, and grants the thing you actually care about.

Braintrust — proprietary engine, your object storage

Braintrust is eval-first: the tracing is in service of regression testing rather than the other way round, and that shows in where its ergonomics are good. Its hybrid deployment, available on enterprise plans, uses Terraform to put the data plane — API, Postgres, Redis, object storage and Brainstore — in your own cloud account while Braintrust hosts the control plane. Brainstore is its own Rust engine, writing spans to object storage and indexing them with the open-source Tantivy library.

So the bytes can be in your account while the thing that can read them efficiently is not yours. That is a genuinely different risk profile from both the open-core and the pure-SaaS cases: your data-residency and sovereignty questions get good answers, and your exit question gets a worse one than the storage location suggests.

LangSmith — managed by default, self-hosted by contract

LangSmith is proprietary, usage-priced, and self-hostable only on the enterprise plan. For a team already writing LangChain and LangGraph, the integration cost is close to zero and the trace views match the abstractions in the code, which is a real and underrated advantage — a trace that mirrors your own control flow gets read, and one that does not, does not. The trade is that on every plan below enterprise, both the code and the store are someone else's, so your historical traces exist at a vendor's pagination limits.

The licence tells you about the code, not about the data

It is worth stating the failure mode directly, because it is the most common mistake in this category. Teams rank these four on how open the licence is, conclude that the MIT option is the portable one and the proprietary ones are traps, and then make a choice that does not survive contact with the actual migration.

Sort by exit cost instead and the ordering changes. Phoenix, under a licence that is not open source by the OSI definition, gives you a database you can dump. Braintrust, proprietary throughout, can put the span bytes in your S3 bucket. LangSmith's enterprise self-host puts the whole thing in your VPC under a commercial agreement. Meanwhile an MIT licence on a hosted plan you never self-host buys you the theoretical right to run code you are not running — which is worth something as insurance against the vendor disappearing, and nothing at all for getting last quarter's traces into a different tool next month.

# Rank on these, in this order

1  can I run a bulk read of historical spans, un-paginated?
2  do the stored attribute names come from my code or the vendor's SDK?
3  where do the bytes physically sit, and under whose contract?
4  what is the licence?

# Note what is NOT on the list

   dashboard quality        — you will look at it twice a week
   number of integrations   — you need two of them
   judge library size       — you will write your own judges

None of which makes the licence worthless: a commercial licence that can be withdrawn is a vendor-risk entry, and the right place for it is third-party and vendor risk, not the portability column.

Own the shim and the choice gets smaller

Three exit paths, and what each one costs when you leave Three columns describing how you get historical traces out of a platform. A paginated REST export is the common case: it works, it is rate-limited, and a year of spans becomes a multi-day job nobody budgets for. A database you already operate means export is a query, so leaving costs a dump and a schema translation. Owning the instrumentation shim is the only path where leaving costs nothing historical, because the attribute names you store were never the vendor's to define. Portability is a property of the write path, not of a feature list PAGINATED REST EXPORT The default everywhere it does work rate-limited per call a year of spans is a multi-day job nobody budgeted in practice: you re-start the history instead A DATABASE YOU OPERATE Langfuse, Phoenix export is a query no rate limit to plan around you still owe a schema translation on the way out cost: a dump plus a mapping, measured in days not quarters A SHIM YOU OWN ~200 lines, your repo you name the attributes the platform is a renderer dual-write during migration is a config change cost of leaving: the backfill only, and you can skip it The third column is compatible with all four platforms, which is the point: it is the decision you make before you choose one.
The third column is compatible with all four platforms, which is why it comes first.

The most useful thing in this post is not a ranking. It is that the portability decision happens upstream of the platform, in roughly two hundred lines of code you write once: a thin layer between your agent loop and whichever SDK you export with, where you decide the attribute names for the things you will still care about in a year — run id, task class, tool name, retry index, token counts split by cache state, judge verdict and version, human override. Emit those under names you chose, map them to the vendor's expectations at the boundary, and the platform becomes a renderer over your schema rather than the owner of it.

Two properties follow, and they are what make the shim worth the afternoon. Dual-writing during a migration becomes a configuration change rather than a project, so you can run two platforms side by side for a month and compare them on your own traffic instead of on a demo dataset. And the churn in the gen_ai.* conventions becomes someone else's problem: when an attribute is renamed upstream, you change one mapping table, and every dashboard and saved query you own keeps working. Pin the convention version you build against and move it deliberately, the same discipline as anything else your vendor can change without a deploy.

Two things the shim does not fix. It does not retroactively port history you already accumulated in a vendor's model, which is the argument for doing this before the first production run rather than after the first renewal. And it does not reduce volume: whatever you choose, trace cost scales with traffic and the lever is sampling, so decide retention and sampling policy at the same time — see trace sampling and retention, and keep redaction in the shim where it belongs rather than in a vendor setting.

When to pick which

SituationPickBecause
Regulated data, no vendor contract in sightLangfuseMIT core self-hosts on every tier; the store is yours without a negotiation.
You need a trace readable todayPhoenixOne process, no event caps, retrieval-shaped span kinds out of the box.
Your bottleneck is regression testing, not debuggingBraintrustEval is the primary object; the hybrid data plane answers residency.
LangChain or LangGraph throughout, small teamLangSmithZero integration cost and traces that mirror your own control flow.
You expect to change your mind within a yearAny, plus the shimThe shim is what makes the first answer reversible.

And the honest summary of the field: none of these four is bad, the dashboards are converging, and the difference that will matter to you in eighteen months is whether the traces you are accumulating right now are written in a vocabulary you control. That is a decision you are making today whether or not you are making it deliberately.

FAQ

If all four accept OTLP, can I just switch later?

You can switch where new spans go in an afternoon. What does not switch is the history: each platform stores spans in its own model, so moving a year of traces means a bulk read through whatever export path exists, plus a schema translation. That is why the question to ask in a trial is about bulk read, not about ingest.

Are the OpenTelemetry GenAI conventions safe to standardise on?

Safe to adopt, not safe to treat as frozen. Every gen_ai.* attribute still carries Development stability, and the conventions now live in a dedicated repository. Instrument against them, pin the version you build on, and keep a mapping layer you own so a rename upstream is a one-file change.

Is Phoenix open source?

Source-available. Elastic License 2.0 permits broad internal use and modification but not offering Phoenix as a hosted managed service, and it is not OSI-approved. For most teams that distinction changes nothing operationally; for a procurement checklist with a binary field, it changes the answer.

Does Braintrust's hybrid deployment mean I own my data?

You own the bytes and their location — the data plane, including Brainstore, runs in your cloud account via Terraform on enterprise plans. The engine that reads them efficiently remains proprietary, so residency is answered well and exit is answered less well than the storage location implies.

What is the one thing to build before choosing?

The instrumentation shim: your own attribute names for run id, task class, tool name, retry index, token counts by cache state, judge verdict and version. Roughly two hundred lines, written once, and it converts the platform choice from a one-way door into a configuration value.

Further reading

On this wiki:

Project sources: