AI Blog

RouteLLM vs Not Diamond vs vLLM Semantic Router vs OpenRouter Auto

OpenRouter’s Auto Router runs Not Diamond underneath, so four products are three routing decisions. The one that matters for agents is not which model — it is how much computation a query deserves, which is what the vLLM Semantic Router classifies. And a router is a classifier whose errors are silent: it returns a valid, slightly worse answer with a 200, so the savings are the only number you will see unless you keep a held-out set. Inside an agent loop, per-step routing fights prompt caching and usually loses.

By Agentic AI Wiki 15 min read

Two of these four are the same router wearing different billing: OpenRouter's Auto Router runs Not Diamond underneath, so a team hedging by adopting both has adopted one. That is the smaller surprise. The larger one is that a router is a classifier you have placed in the request path, and its errors are silent — a misroute returns a valid, slightly worse answer with a 200 status — so adopting one obliges you to run a quality evaluation you probably do not have. And inside an agent loop, per-step routing collides with prompt caching, which usually saves more than the routing does.

At a glance

Four products, three distinct routing decisions, and two completely different places to stand.

ProjectOriginWhat it decidesWhere it runs
RouteLLM LMSYS / Berkeley, open source, ICLR 2025 Strong model or weak model, on a tunable threshold A library in your process
Not Diamond Commercial, seed-funded, SaaS Which model, from a learned per-query preference A hosted call before your provider call
OpenRouter Auto Router OpenRouter, beta, free to use Same decision — it runs Not Diamond Inside the gateway you already call
vLLM Semantic Router vLLM ecosystem, Apache 2.0 Intent and complexity — and whether to reason at all An Envoy filter in your own serving stack
Feature matrix for four model routers Rows are RouteLLM, Not Diamond, OpenRouter Auto Router and vLLM Semantic Router. Columns are self-hosted, ships an eval harness, breadth across external providers, control over reasoning effort, and whether an extra network hop is added. RouteLLM is strong on self-hosting and evaluation, weak on provider breadth. Not Diamond and Auto Router are strong on breadth and weak on self-hosting and evaluation; Auto Router adds no extra hop because it lives in the gateway. vLLM Semantic Router is strong on self-hosting and reasoning control and weak on external providers. Where each router is strong Self-hosted Eval harness Provider breadth Effort control Extra hop RouteLLM Strong Strong Weak — a pair Weak None Not Diamond Weak — hosted Weak Strong Weak One round trip Auto Router Weak — hosted Weak Strong Weak In the gateway Semantic Router Strong Medium Weak — local Strong In-path filter Strong Medium Weak No row is strong everywhere, and the two hosted rows share one decision engine. Effort control is the column that matters most on reasoning-model traffic, and only one row has it. Provider breadth is the column that matters least if you self-host open weights.
Read the last column first: three of the four add a decision hop before the model call.

The categories matter more than the cells. RouteLLM and vLLM Semantic Router are things you run; Not Diamond and Auto Router are things you call. That single difference decides your latency floor, your failure modes, and whether the routing decision is auditable after the fact.

Where each one sits

Three placements for a model router on the request path A request leaves the application and reaches a model. Three placements are shown. In-process: a RouteLLM classifier runs inside the application and picks a strong or weak model before any network call is made. At the proxy: an Envoy external processor in your own serving stack classifies intent and complexity, may disable reasoning mode, and forwards to a local backend. Hosted: the application calls a routing service, or calls a gateway whose auto endpoint calls that same service, and a model choice comes back before the provider call. A note records that OpenRouter Auto and Not Diamond resolve to the same decision engine. Three places the routing decision can live In your process RouteLLM Strong or weak, on a threshold you tune. Provider call No extra hop. Decision is local, logged, and yours to retrain. At your proxy vLLM Semantic Router Envoy ext_proc, ModernBERT intent and complexity. Local backend Reasoning on or off. Hosted Your application Asks first, then calls. Not Diamond Per-query model preference. Provider call Plus one network round trip. OpenRouter Auto Router Same decision, inside the gateway you already call. The proxy row also carries a semantic cache, a PII check and a prompt guard. Auto Router and Not Diamond resolve to one engine: adopting both is adopting one. The top two rows keep the decision inside your boundary, where it can be logged and replayed.
Three placements, and one of the three boxes is shared by two of the four products.

RouteLLM — the honest place to start

RouteLLM came out of LMSYS, the Chatbot Arena team, and was published at ICLR 2025. It frames routing as a binary: send this query to the strong model or the weak one, with a threshold you tune to a cost-quality point you choose. It ships several router implementations — similarity-weighted ranking, a matrix factorisation model, a BERT classifier — and, more importantly, the eval harness they were measured with. The headline result from the paper is roughly GPT-4-level quality on most of a benchmark at around a quarter of the cost.

Treat that number as a method demonstration rather than a forecast. What RouteLLM actually gives you is the only honest artefact in this category: a way to measure the thing on your own queries before you believe anyone's percentage. The cost is that it is a research codebase — you own the model pair, the threshold, the retraining, and the plumbing.

Not Diamond, and the Auto Router that is Not Diamond

Not Diamond is a hosted routing layer: you send it the query, it returns which model should handle it, and you make the provider call yourself. It is explicitly not a gateway — it sits on top of OpenRouter, Hugging Face, or whatever provider stack you already have. The company raised a small seed round with a notably credentialed angel list, and its routing is learned per query rather than thresholded on a single strong/weak axis.

OpenRouter's Auto Router is the same decision, delivered differently. The openrouter/auto endpoint is in beta, costs nothing beyond the underlying model's tokens, picks from a catalogue spanning hundreds of models across dozens of providers — and runs Not Diamond as the engine. For most teams that packaging is strictly better: no extra hop, no second vendor, no separate key. It is worth knowing the lineage anyway, because "we evaluated two routers" is not the diversification it sounds like, and because a dependency you did not choose is still a dependency.

vLLM Semantic Router — the one that is not really a model router

This is the odd entry and, for agent workloads, often the most relevant. It runs as an Envoy external processor in front of your own vLLM backends: Envoy intercepts the request, hands it over gRPC to a classifier, and the classifier decides where it goes. The core is Rust on Hugging Face Candle wrapped in Go for the ext_proc interface, and the classifier is a ModernBERT model scoring intent and complexity. It is Apache 2.0, and it ships semantic caching, PII detection and a prompt guard alongside the routing.

The decision it makes is the interesting part. Rather than only choosing between models, it decides whether reasoning is worth enabling for this query — auto reasoning-mode adjustment. The project's published MMLU-Pro figures with a Qwen3 30B model are a 10.2% accuracy improvement alongside a 47.1% latency reduction and 48.5% fewer tokens. Accuracy going up while tokens go down is the tell: on easy queries, reasoning is not merely expensive, it is sometimes actively worse. That is the same argument as adaptive thinking and effort budgets, implemented at the proxy.

What a router actually decides

The three questions these routers actually ask Three columns. Is this hard: a tunable strong-versus-weak threshold, one dial you can explain, but you must choose the model pair yourself. Which model suits this query: a learned preference across a broad catalogue, wide coverage, with a decision you cannot reconstruct afterwards. How much computation does this deserve: an effort and intent classifier that can disable reasoning entirely, which is the axis that dominates spend on reasoning-model traffic. Three routers, three different questions "Is this hard?" RouteLLM A tunable strong/weak threshold. One dial. You pick the model pair, and you own retraining. Explainable to finance. "Which model suits this?" Not Diamond / Auto Router A learned preference over a broad catalogue. Coverage bought at the price of an opaque call. Hard to reconstruct later. "How much compute?" vLLM Semantic Router Intent and complexity — and reasoning on or off. On reasoning traffic this spread beats vendor price. Accuracy can rise as tokens fall. The first two compete on the same axis. The third competes on the axis agent bills are actually built from. Pick the question before the product: a uniform-difficulty workload needs none of the three.
Three different questions. Only the third one is about the axis agent spend actually moves on.

Compare the three decisions on the same axis and the split is not about accuracy, it is about what question is being asked. A strong/weak threshold asks "is this hard?" and gives you one dial you can explain to a finance team. A learned per-query preference asks "which of these models tends to do best on queries like this?" and gives you breadth at the price of a decision you cannot reconstruct. An effort classifier asks "how much computation does this deserve?" — and on reasoning-model traffic that is where the money is, because the spread between a short answer and a long chain of thought dwarfs the spread between two vendors' per-token prices.

The published savings ranges — commonly quoted between 30% and 85% — are all real and all measured on someone else's traffic. A router's return is a property of your query mix: it pays when a large head of your requests are genuinely easy and a small tail is genuinely hard. Chat assistants look like that. Extraction pipelines do not, because every request is the same difficulty and the right answer is to pick one small model and stop paying a classifier to re-discover that. Measure your own distribution before you believe a percentage; unit economics is the frame, and cost per successfully completed task is the number.

The evaluation you now owe

This is the part that gets skipped, and it is the reason router projects quietly stall six months in. Every other component in your stack fails loudly. A router does not: when it sends a hard query to the small model, you get a 200, a fluent answer, and a slightly worse outcome. There is no error rate to alert on, no latency spike, no log line that says this was the wrong choice.

  • Router quality is invisible in ops metrics and visible only against a held-out set with known-good answers. If you do not have one, the router is unfalsifiable and its savings are the only number you will ever see — which is precisely the failure shape that makes it feel like a success.
  • Every model release moves the decision boundary. A learned router trained on last quarter's model pair is making calibrated decisions about capabilities that have changed. Routers are perishable in a way that gateways are not, and quality regression detection is the thing that tells you the fruit has turned.
  • Log the decision, not just the outcome. Store which model was chosen, by which router version, with what score, on every request. Without that field you cannot answer "did quality drop because of the router or because of the model?", which is the first question anyone will ask.
  • Shadow before you switch. Run the router in parallel on a slice of traffic, send everything to the strong model anyway, and compare what the router would have chosen against what the strong model actually produced. The cost of that experiment is a fraction of the savings, and it is the only way to size the regression in advance.
  • Budget the hop. A hosted router adds a network round trip before the model call — typically tens of milliseconds, occasionally much worse — and a dependency that can fail. Decide now what happens when the router is down: defaulting to the strong model is right, defaulting to the cheap one is a silent quality incident, and graceful degradation is where that choice belongs.

Routing inside an agent loop breaks the cache

Everything above applies to chat traffic. Agents change the arithmetic, and not in the router's favour.

An agent step re-sends the whole transcript: system prompt, tool definitions, every prior tool result. That prefix is stable and large, which is exactly what prompt caching is for, and cached input reads are typically discounted by around 90% against fresh input. Caches are per-model. Route step 4 to a different model than step 3 and the prefix is cold again — you pay full price for the entire accumulated transcript in order to save a fraction on one step's tokens. In most agent loops the cache discount is the larger number, which means per-step routing can cost you money while every dashboard says it is saving some. See prompt caching and context caching economics for the arithmetic.

There is a correctness argument too, and it is the stronger one. Models differ in how they emit tool calls, how eagerly they call them, how they format structured output, and how they behave when a tool returns an error. Swapping models mid-trajectory means the run's behaviour changes halfway through for reasons unrelated to the task, and reproducing a failure requires knowing which model handled which step. That is a real debugging tax for a saving that the cache has usually already eaten.

The placement that does work is the task boundary rather than the step boundary: classify the incoming task once, pick a model for the whole run, and keep the prefix warm for its duration. That preserves the cache, keeps the trajectory coherent, and still captures most of the spread — because task difficulty, not step difficulty, is what determines whether you needed the frontier model at all. It is also the placement the router pattern describes, and it is why the effort-classifier shape generalises better here than the model-picker shape does.

When to pick which

If your situation is…RouteLLMNot Diamond / Auto RoutervLLM Semantic Router
You want to measure before you commitYes — it ships the harnessHard; the decision is opaqueYes, on your own stack
Already on a gateway, want the easy winNoThis, and no extra vendorNo
Self-hosted open weightsLibrary onlyNot the shapeThis is the case
Reasoning models dominate the billIndirectIndirectDirect — effort control
Agent loops with long transcriptsAt task boundary onlyAt task boundary onlyAt task boundary only
Uniform-difficulty pipelineSkip itSkip itSkip it

The last row is the one most teams belong in and least want to hear. If every request in a workload is the same shape, the correct router is a decision you make once, in code, for free.

FAQ

Is OpenRouter's Auto Router really the same as Not Diamond?

The Auto Router endpoint uses Not Diamond as its routing engine, so the decision comes from the same place. What differs is packaging: Auto Router arrives inside a gateway you are already calling, with no extra hop, no second contract and no separate key. If you were choosing between them for redundancy, you were choosing between one engine and the same engine.

Do these routers actually cut costs by 30–85%?

Those figures are real measurements on traffic that had a fat head of easy queries and a thin tail of hard ones. Your saving is a property of your own difficulty distribution, not of the router. Measure it on a held-out slice of your own requests before you plan against a number from someone else's workload.

Can I put a router inside an agent loop?

At the task boundary, yes. At the step boundary it fights prompt caching — a model switch makes the accumulated transcript a cache miss, and the discount on cached reads is usually larger than the saving on one step — and it makes tool-calling behaviour change mid-trajectory, which is a debugging cost you pay forever.

What is different about the vLLM Semantic Router?

It runs as an Envoy external processor in your own serving stack rather than as a hosted call, and its primary decision is how much computation a query deserves — including whether to enable reasoning at all — rather than which vendor's model to use. On reasoning-heavy traffic that axis moves far more money than model choice does.

What breaks first when a router goes wrong?

Nothing, visibly, which is the problem. A misroute returns a valid response with a normal status code and slightly worse content. Without a held-out evaluation set and the chosen-model field logged on every request, router regressions are indistinguishable from ordinary model variance.

Further reading

On this wiki:

Project sources: