AI Blog

OTel GenAI vs OpenInference vs OpenLLMetry vs OpenLIT: the neutral option is the one still moving

Of the four ways to shape an agent trace, the only one that calls itself the standard is the only one you cannot pin: on 12 June 2026 OpenTelemetry deprecated all sixty gen_ai attributes and moved them to a repository that still has no tagged release and a TODO where its schema URL should be. Eight renames and two deletions landed in that one version — including both token-usage attributes. Pick by the vocabulary your backend dispatches on, translate at the collector, and never point a cost chart at a Development-stability attribute name.

By Agentic AI Wiki 12 min read

Of the four ways to shape an agent trace, exactly one calls itself the standard, and it is the only one you cannot pin to a version. On 12 June 2026 OpenTelemetry's v1.42.0 deprecated the entire gen_ai namespace in core semantic conventions and moved it to a dedicated repository — which today has 653 commits, no tagged release, and a README whose schema-URL section still reads TODO. Every gen_ai.* attribute, span, metric and event carries the stability badge Development; none is Stable. So "we instrument to the open standard" is not the safe answer it sounds like, and the thing that actually decides whether your dashboards survive next quarter is which span vocabulary your backend keys its features off.

At a glance

Two of these are specifications, two are SDKs, and the distinction is where most of the confusion starts.

ProjectShipped byWhat it actually isCan you pin it?
OTel GenAI conventionsOpenTelemetry GenAI SIGAttribute, span, metric and event names. No SDK of its own.No tagged release; commit or snapshot only.
OpenInferenceArize AIA spec plus instrumentors for Python, JS, Java and Go.Yes — versioned packages, stable span kinds.
OpenLLMetryTraceloopAn SDK with its own span shape, partly upstreamed to OTel.Yes — versioned, but the shape is mid-migration.
OpenLITOpenLIT projectOTel-native SDK plus a platform, broadest scope.Yes — versioned, tracks gen_ai.* as it moves.

Rough community size, for context rather than ranking: OpenLLMetry around 7.5k GitHub stars, OpenLIT around 2.8k, OpenInference around 1.3k, and the OTel GenAI conventions repository around 414 — a number that mostly reflects that specifications do not get starred the way libraries do.

Where each GenAI tracing convention leans hardest A four-by-four matrix. Rows are the OpenTelemetry GenAI conventions, OpenInference, OpenLLMetry and OpenLIT. Columns are pinnable release, agent-shaped span vocabulary, auto-instrumentation breadth and backend portability. The OTel GenAI row is weak on pinnable release and medium on span vocabulary while strong on portability; OpenInference is strong on span vocabulary; OpenLLMetry is strong on instrumentation breadth and medium elsewhere; OpenLIT is strong on portability and medium on vocabulary. FOUR CONVENTIONS, FOUR AXES PINNABLE? AGENT SPANS BREADTH PORTABILITY OTel GenAI conventions Weak (no tag) Medium Weak (no SDK) Strong OpenInference Strong Strong (11 kinds) Medium Medium OpenLLMetry Medium Weak (straddles) Strong Medium OpenLIT Medium Medium Medium (+GPU) Strong Weak Medium Strong
No row is strong everywhere, and the column teams weight least is the one that bit them.

The neutral option is the one still moving

What the June 2026 gen_ai deprecation did to sixty attributes Three columns covering the sixty gen_ai attributes deprecated in OpenTelemetry semantic conventions v1.42.0 on 12 June 2026. Fifty moved to a new repository keeping the same attribute string. Eight were renamed, including both token-usage attributes and gen_ai.system becoming gen_ai.provider.name. Two, gen_ai.prompt and gen_ai.completion, were removed with no replacement, with content capture moving to opt-in message attributes. SIXTY gen_ai ATTRIBUTES DEPRECATED IN v1.42.0, 12 JUNE 2026 50 moved, same string Informational only new repository, identical attribute name 8 renamed Charts under-count silently usage.prompt_tokens → usage.input_tokens 2 removed outright No rename to apply gen_ai.prompt, gen_ai.completion
Fifty of the sixty attributes kept their string. The ten that did not include every one your cost dashboard reads.

Start with what happened in June, because it reframes the whole comparison. Semantic conventions v1.42.0 deprecated all sixty gen_ai attributes in the core repository and relocated the namespace — along with the MCP conventions — into open-telemetry/semantic-conventions-genai, explicitly so that this area could iterate faster than core's stability bar permits. Of the sixty, fifty kept their exact string and simply changed address, eight were renamed, and two were removed with nothing taking their place.

The ten that moved are the ones that matter. gen_ai.usage.prompt_tokens became gen_ai.usage.input_tokens and gen_ai.usage.completion_tokens became gen_ai.usage.output_tokens, so every cost dashboard built on the old names under-counts silently the moment an emitter upgrades — no error, no gap in the chart, just a smaller number. gen_ai.system became gen_ai.provider.name, which breaks per-provider breakdowns the same way. And gen_ai.prompt and gen_ai.completion were deleted rather than renamed: content capture now goes through the opt-in gen_ai.input.messages and gen_ai.output.messages, plus gen_ai.system_instructions. If your replay tooling reads prompts off spans, that is not a rename you can sed.

None of this is mismanagement. A fast-moving area was given a faster-moving home, which is the correct call. But it does invert the usual argument for vendor-neutral instrumentation. The pitch is that neutrality saves you a migration; the reality in 2026 is that the neutral layer is the one issuing the migrations, while a vendor SDK that froze its shape two years ago has issued none. Neutrality still wins on a long horizon. It is not free on a one-year one, and the cost lands on whoever wrote queries against attribute names.

Span kinds are the contract; attribute names are the churn

Where a convention choice lands in the tracing stack Four stacked layers from application code through the instrumentation SDK to the emitted span shape and finally the backend features that consume it. The span-kind vocabulary inside the span shape is marked as the contract backend features dispatch on; attribute names are marked as the layer that changes between versions. A collector processor sits beside the span-shape layer as the place to absorb renames. THE STACK WHAT MOVES Application code framework calls, your agent loop Should never name a convention Instrumentation SDK OpenInference · OpenLLMetry · OpenLIT Upgrades arrive with bug fixes Emitted span shape span kind — the contract the backend dispatches on attribute names — the layer that gets renamed Span kind: stable, load-bearing Attributes: 10 of 60 changed in one release Backend features agent view, eval pipeline, replay, cost charts Keyed on span kind, queried on attributes
Your backend's evals, replay and agent views read the span kind. Rename an attribute and a chart breaks; change the span kind and the product stops working.

Here is the distinction almost every comparison misses. An attribute name is a label on a span; if it changes, a query breaks and you fix the query. A span kind is the thing a backend dispatches on — whether this span is rendered as an LLM call, a retrieval step, a tool invocation or an agent turn, whether it enters the eval pipeline, whether it can be replayed. Break that and the feature does not degrade, it disappears.

The four projects answer this very differently. OpenInference makes the span kind explicit and required: a single openinference.span.kind attribute with eleven values — LLM, EMBEDDING, CHAIN, RETRIEVER, RERANKER, TOOL, AGENT, GUARDRAIL, EVALUATOR, PROMPT and DECISION — alongside namespaces for the payloads (llm.*, document.*, message.*, tool.*, evaluation.*, session.*). That vocabulary is the most agent-shaped of the four, and it is also the oldest unchanged thing in this comparison, because Arize froze it early to build a product on it.

OTel GenAI gets there by operation name rather than a span-kind enum, and the agent-relevant work is recent: v1.41.0 added invoke_workflow as an operation and split invoke_agent into client and internal spans, alongside the first streaming metrics (gen_ai.client.operation.time_to_first_chunk and .time_per_output_chunk). That split is the right modelling decision and it is also a breaking change in how agent spans nest, which is exactly the kind of movement you do not want underneath a product feature.

OpenLLMetry sits in the awkward middle by history rather than by choice. Traceloop shipped its span shape before the GenAI conventions existed, then upstreamed its conventions into OpenTelemetry — the project's own README says so — and the result is an SDK whose output straddles two vocabularies and has not fully converged. Its issue tracker carries the expected report that gen_ai.prompt and gen_ai.completion are now deprecated upstream. OpenLIT took the other route: OTel-native from the start, following gen_ai.* as it moves, with the widest surface of the four — it traces LLM calls, tool calls, retrieval, embeddings, vector-store operations, MCP requests and sub-agent activity, and on coding agents it also spans file reads and edits, shell commands and searches, correlated with GPU utilisation from NVIDIA, AMD and Intel collectors.

What breadth buys and what it costs

Auto-instrumentation coverage is the axis teams actually feel in week one, and the three SDKs are genuinely different here. OpenInference ships instrumentors across Python, JavaScript, Java and Go for a long list of providers and frameworks — OpenAI, Anthropic, Bedrock, VertexAI, Mistral, Groq, LangChain, LlamaIndex, CrewAI, Haystack, PydanticAI among them — which is unsurprising, since every framework it covers is a framework Phoenix needs to render. OpenLLMetry's strength is the same breadth on the provider and vector-store side, Python first with a separate JS/TS distribution, and the deepest set of third-party backend integrations of the four. OpenLIT's strength is scope rather than language count: the GPU correlation and the coding-agent spans are things neither of the others attempts.

The cost of breadth is that auto-instrumentation is where span shape leaks into your data without a decision. A library that patches forty clients emits whatever its authors decided each one maps to, and when the convention underneath moves, the mapping moves with it on an upgrade you took for a bug fix. This is the practical reason to put a transformation at the collector rather than trusting the SDK: a processor that renames attributes and normalises span kinds on the way out is a config change you ship once, and it keeps the churn out of forty call sites. The alternative — decorating application code with one vendor's helpers — is the lock-in that instrumenting the wire format rather than the vendor is about avoiding.

One more asymmetry worth naming. Backends are not symmetric about which vocabulary they prefer: Phoenix is OpenInference-native, Langfuse and Braintrust accept OTLP and map what they can, and translation is lossy in the direction of the richer vocabulary. Converting OpenInference spans into gen_ai.* loses the span-kind distinctions that have no counterpart; converting the other way invents them. If your eval or annotation workflow depends on a distinction that only one vocabulary expresses, that is not a mapping problem, it is a product-feature problem.

When to pick which

SituationPickWhy
You run Phoenix, or you want agent-shaped spans todayOpenInferenceEleven explicit span kinds, stable for years, and the backend reads them natively.
You are standardising across many services and vendorsOTel GenAI conventionsIt is where everything converges — but pin a commit, not "latest", and expect renames.
You need the widest provider and vector-store coverage, Python firstOpenLLMetryDeepest integration surface; accept that the span shape is mid-convergence.
You self-host inference, or you operate coding agentsOpenLITGPU correlation and shell/file/search spans exist nowhere else in this set.
You are building a cost or usage dashboardNone of them, directlyDerive your own metric from the spans and own the rename. Do not query gen_ai.* from a chart.

The last row is the recommendation that matters most and the one nobody puts in a comparison table. A dashboard that reads gen_ai.usage.input_tokens directly is a query against a Development-stability name in a repository with no tagged release. Compute your token and cost metrics once, in a collector processor or a small job, emit them under names you control, and point every chart at those. It is an afternoon of work and it is the difference between a rename being a config change and a rename being a quarter of silently wrong finance numbers.

FAQ

Are the OpenTelemetry GenAI conventions stable yet?

No. Every gen_ai.* attribute, span, metric and event in the registry carries the Development stability badge, and the dedicated semantic-conventions-genai repository has no tagged release and no finalised schema URL. You can pin a commit or a dated snapshot; you cannot pin a version.

Did OpenLLMetry's conventions become the OpenTelemetry ones?

Partly. Traceloop upstreamed its semantic conventions into OpenTelemetry, so the GenAI conventions reflect that contribution. But the OpenLLMetry SDK shipped its own span shape before those conventions existed and has not fully converged on them, so in practice you are emitting something between the two.

What actually broke in the June 2026 deprecation?

Eight renames and two deletions out of sixty attributes. The renames include both token-usage attributes and gen_ai.system becoming gen_ai.provider.name; the deletions are gen_ai.prompt and gen_ai.completion, replaced by the opt-in message attributes. The other fifty emit the same string from a new home.

Can I migrate between these later without touching application code?

Mostly, if you kept instrumentation generic and put the mapping in a collector processor. Attribute renames translate cleanly. Span-kind distinctions do not: a vocabulary that never expressed GUARDRAIL or EVALUATOR spans cannot have them translated into it, so backend features keyed on those will not appear.

Which one should a small team default to?

Whichever vocabulary the backend you have chosen reads natively, with one collector-side processor between you and it. Picking the instrumentation before the consumer is how teams end up emitting a vocabulary nothing they run can use.

Further reading

On this wiki:

Project sources: