Of the four ways to shape an agent trace, exactly one calls itself the standard, and it is the only one you cannot pin to a version. On 12 June 2026 OpenTelemetry's v1.42.0 deprecated the entire gen_ai namespace in core semantic conventions and moved it to a dedicated repository — which today has 653 commits, no tagged release, and a README whose schema-URL section still reads TODO. Every gen_ai.* attribute, span, metric and event carries the stability badge Development; none is Stable. So "we instrument to the open standard" is not the safe answer it sounds like, and the thing that actually decides whether your dashboards survive next quarter is which span vocabulary your backend keys its features off.
At a glance
Two of these are specifications, two are SDKs, and the distinction is where most of the confusion starts.
| Project | Shipped by | What it actually is | Can you pin it? |
|---|---|---|---|
| OTel GenAI conventions | OpenTelemetry GenAI SIG | Attribute, span, metric and event names. No SDK of its own. | No tagged release; commit or snapshot only. |
| OpenInference | Arize AI | A spec plus instrumentors for Python, JS, Java and Go. | Yes — versioned packages, stable span kinds. |
| OpenLLMetry | Traceloop | An SDK with its own span shape, partly upstreamed to OTel. | Yes — versioned, but the shape is mid-migration. |
| OpenLIT | OpenLIT project | OTel-native SDK plus a platform, broadest scope. | Yes — versioned, tracks gen_ai.* as it moves. |
Rough community size, for context rather than ranking: OpenLLMetry around 7.5k GitHub stars, OpenLIT around 2.8k, OpenInference around 1.3k, and the OTel GenAI conventions repository around 414 — a number that mostly reflects that specifications do not get starred the way libraries do.
The neutral option is the one still moving
Start with what happened in June, because it reframes the whole comparison. Semantic conventions v1.42.0 deprecated all sixty gen_ai attributes in the core repository and relocated the namespace — along with the MCP conventions — into open-telemetry/semantic-conventions-genai, explicitly so that this area could iterate faster than core's stability bar permits. Of the sixty, fifty kept their exact string and simply changed address, eight were renamed, and two were removed with nothing taking their place.
The ten that moved are the ones that matter. gen_ai.usage.prompt_tokens became gen_ai.usage.input_tokens and gen_ai.usage.completion_tokens became gen_ai.usage.output_tokens, so every cost dashboard built on the old names under-counts silently the moment an emitter upgrades — no error, no gap in the chart, just a smaller number. gen_ai.system became gen_ai.provider.name, which breaks per-provider breakdowns the same way. And gen_ai.prompt and gen_ai.completion were deleted rather than renamed: content capture now goes through the opt-in gen_ai.input.messages and gen_ai.output.messages, plus gen_ai.system_instructions. If your replay tooling reads prompts off spans, that is not a rename you can sed.
None of this is mismanagement. A fast-moving area was given a faster-moving home, which is the correct call. But it does invert the usual argument for vendor-neutral instrumentation. The pitch is that neutrality saves you a migration; the reality in 2026 is that the neutral layer is the one issuing the migrations, while a vendor SDK that froze its shape two years ago has issued none. Neutrality still wins on a long horizon. It is not free on a one-year one, and the cost lands on whoever wrote queries against attribute names.
Span kinds are the contract; attribute names are the churn
Here is the distinction almost every comparison misses. An attribute name is a label on a span; if it changes, a query breaks and you fix the query. A span kind is the thing a backend dispatches on — whether this span is rendered as an LLM call, a retrieval step, a tool invocation or an agent turn, whether it enters the eval pipeline, whether it can be replayed. Break that and the feature does not degrade, it disappears.
The four projects answer this very differently. OpenInference makes the span kind explicit and required: a single openinference.span.kind attribute with eleven values — LLM, EMBEDDING, CHAIN, RETRIEVER, RERANKER, TOOL, AGENT, GUARDRAIL, EVALUATOR, PROMPT and DECISION — alongside namespaces for the payloads (llm.*, document.*, message.*, tool.*, evaluation.*, session.*). That vocabulary is the most agent-shaped of the four, and it is also the oldest unchanged thing in this comparison, because Arize froze it early to build a product on it.
OTel GenAI gets there by operation name rather than a span-kind enum, and the agent-relevant work is recent: v1.41.0 added invoke_workflow as an operation and split invoke_agent into client and internal spans, alongside the first streaming metrics (gen_ai.client.operation.time_to_first_chunk and .time_per_output_chunk). That split is the right modelling decision and it is also a breaking change in how agent spans nest, which is exactly the kind of movement you do not want underneath a product feature.
OpenLLMetry sits in the awkward middle by history rather than by choice. Traceloop shipped its span shape before the GenAI conventions existed, then upstreamed its conventions into OpenTelemetry — the project's own README says so — and the result is an SDK whose output straddles two vocabularies and has not fully converged. Its issue tracker carries the expected report that gen_ai.prompt and gen_ai.completion are now deprecated upstream. OpenLIT took the other route: OTel-native from the start, following gen_ai.* as it moves, with the widest surface of the four — it traces LLM calls, tool calls, retrieval, embeddings, vector-store operations, MCP requests and sub-agent activity, and on coding agents it also spans file reads and edits, shell commands and searches, correlated with GPU utilisation from NVIDIA, AMD and Intel collectors.
What breadth buys and what it costs
Auto-instrumentation coverage is the axis teams actually feel in week one, and the three SDKs are genuinely different here. OpenInference ships instrumentors across Python, JavaScript, Java and Go for a long list of providers and frameworks — OpenAI, Anthropic, Bedrock, VertexAI, Mistral, Groq, LangChain, LlamaIndex, CrewAI, Haystack, PydanticAI among them — which is unsurprising, since every framework it covers is a framework Phoenix needs to render. OpenLLMetry's strength is the same breadth on the provider and vector-store side, Python first with a separate JS/TS distribution, and the deepest set of third-party backend integrations of the four. OpenLIT's strength is scope rather than language count: the GPU correlation and the coding-agent spans are things neither of the others attempts.
The cost of breadth is that auto-instrumentation is where span shape leaks into your data without a decision. A library that patches forty clients emits whatever its authors decided each one maps to, and when the convention underneath moves, the mapping moves with it on an upgrade you took for a bug fix. This is the practical reason to put a transformation at the collector rather than trusting the SDK: a processor that renames attributes and normalises span kinds on the way out is a config change you ship once, and it keeps the churn out of forty call sites. The alternative — decorating application code with one vendor's helpers — is the lock-in that instrumenting the wire format rather than the vendor is about avoiding.
One more asymmetry worth naming. Backends are not symmetric about which vocabulary they prefer: Phoenix is OpenInference-native, Langfuse and Braintrust accept OTLP and map what they can, and translation is lossy in the direction of the richer vocabulary. Converting OpenInference spans into gen_ai.* loses the span-kind distinctions that have no counterpart; converting the other way invents them. If your eval or annotation workflow depends on a distinction that only one vocabulary expresses, that is not a mapping problem, it is a product-feature problem.
When to pick which
| Situation | Pick | Why |
|---|---|---|
| You run Phoenix, or you want agent-shaped spans today | OpenInference | Eleven explicit span kinds, stable for years, and the backend reads them natively. |
| You are standardising across many services and vendors | OTel GenAI conventions | It is where everything converges — but pin a commit, not "latest", and expect renames. |
| You need the widest provider and vector-store coverage, Python first | OpenLLMetry | Deepest integration surface; accept that the span shape is mid-convergence. |
| You self-host inference, or you operate coding agents | OpenLIT | GPU correlation and shell/file/search spans exist nowhere else in this set. |
| You are building a cost or usage dashboard | None of them, directly | Derive your own metric from the spans and own the rename. Do not query gen_ai.* from a chart. |
The last row is the recommendation that matters most and the one nobody puts in a comparison table. A dashboard that reads gen_ai.usage.input_tokens directly is a query against a Development-stability name in a repository with no tagged release. Compute your token and cost metrics once, in a collector processor or a small job, emit them under names you control, and point every chart at those. It is an afternoon of work and it is the difference between a rename being a config change and a rename being a quarter of silently wrong finance numbers.
FAQ
Are the OpenTelemetry GenAI conventions stable yet?
No. Every gen_ai.* attribute, span, metric and event in the registry carries the Development stability badge, and the dedicated semantic-conventions-genai repository has no tagged release and no finalised schema URL. You can pin a commit or a dated snapshot; you cannot pin a version.
Did OpenLLMetry's conventions become the OpenTelemetry ones?
Partly. Traceloop upstreamed its semantic conventions into OpenTelemetry, so the GenAI conventions reflect that contribution. But the OpenLLMetry SDK shipped its own span shape before those conventions existed and has not fully converged on them, so in practice you are emitting something between the two.
What actually broke in the June 2026 deprecation?
Eight renames and two deletions out of sixty attributes. The renames include both token-usage attributes and gen_ai.system becoming gen_ai.provider.name; the deletions are gen_ai.prompt and gen_ai.completion, replaced by the opt-in message attributes. The other fifty emit the same string from a new home.
Can I migrate between these later without touching application code?
Mostly, if you kept instrumentation generic and put the mapping in a collector processor. Attribute renames translate cleanly. Span-kind distinctions do not: a vocabulary that never expressed GUARDRAIL or EVALUATOR spans cannot have them translated into it, so backend features keyed on those will not appear.
Which one should a small team default to?
Whichever vocabulary the backend you have chosen reads natively, with one collector-side processor between you and it. Picking the instrumentation before the consumer is how teams end up emitting a vocabulary nothing they run can use.
Further reading
On this wiki:
- OpenTelemetry GenAI conventions — why the data model, not the dashboard, is the irreversible choice.
- Tracing and observability for agents — what a useful agent span contains.
- LangSmith vs Langfuse vs Braintrust vs Phoenix — the backend comparison this one sits underneath.
- Trace sampling and retention — the cost side of emitting everything.
- Protocol revisions and deprecation windows — how to absorb a convention that moves under you.