Two of these four are the same router wearing different billing: OpenRouter's Auto Router runs Not Diamond underneath, so a team hedging by adopting both has adopted one. That is the smaller surprise. The larger one is that a router is a classifier you have placed in the request path, and its errors are silent — a misroute returns a valid, slightly worse answer with a 200 status — so adopting one obliges you to run a quality evaluation you probably do not have. And inside an agent loop, per-step routing collides with prompt caching, which usually saves more than the routing does.
At a glance
Four products, three distinct routing decisions, and two completely different places to stand.
| Project | Origin | What it decides | Where it runs |
|---|---|---|---|
| RouteLLM | LMSYS / Berkeley, open source, ICLR 2025 | Strong model or weak model, on a tunable threshold | A library in your process |
| Not Diamond | Commercial, seed-funded, SaaS | Which model, from a learned per-query preference | A hosted call before your provider call |
| OpenRouter Auto Router | OpenRouter, beta, free to use | Same decision — it runs Not Diamond | Inside the gateway you already call |
| vLLM Semantic Router | vLLM ecosystem, Apache 2.0 | Intent and complexity — and whether to reason at all | An Envoy filter in your own serving stack |
The categories matter more than the cells. RouteLLM and vLLM Semantic Router are things you run; Not Diamond and Auto Router are things you call. That single difference decides your latency floor, your failure modes, and whether the routing decision is auditable after the fact.
Where each one sits
RouteLLM — the honest place to start
RouteLLM came out of LMSYS, the Chatbot Arena team, and was published at ICLR 2025. It frames routing as a binary: send this query to the strong model or the weak one, with a threshold you tune to a cost-quality point you choose. It ships several router implementations — similarity-weighted ranking, a matrix factorisation model, a BERT classifier — and, more importantly, the eval harness they were measured with. The headline result from the paper is roughly GPT-4-level quality on most of a benchmark at around a quarter of the cost.
Treat that number as a method demonstration rather than a forecast. What RouteLLM actually gives you is the only honest artefact in this category: a way to measure the thing on your own queries before you believe anyone's percentage. The cost is that it is a research codebase — you own the model pair, the threshold, the retraining, and the plumbing.
Not Diamond, and the Auto Router that is Not Diamond
Not Diamond is a hosted routing layer: you send it the query, it returns which model should handle it, and you make the provider call yourself. It is explicitly not a gateway — it sits on top of OpenRouter, Hugging Face, or whatever provider stack you already have. The company raised a small seed round with a notably credentialed angel list, and its routing is learned per query rather than thresholded on a single strong/weak axis.
OpenRouter's Auto Router is the same decision, delivered differently. The openrouter/auto endpoint is in beta, costs nothing beyond the underlying model's tokens, picks from a catalogue spanning hundreds of models across dozens of providers — and runs Not Diamond as the engine. For most teams that packaging is strictly better: no extra hop, no second vendor, no separate key. It is worth knowing the lineage anyway, because "we evaluated two routers" is not the diversification it sounds like, and because a dependency you did not choose is still a dependency.
vLLM Semantic Router — the one that is not really a model router
This is the odd entry and, for agent workloads, often the most relevant. It runs as an Envoy external processor in front of your own vLLM backends: Envoy intercepts the request, hands it over gRPC to a classifier, and the classifier decides where it goes. The core is Rust on Hugging Face Candle wrapped in Go for the ext_proc interface, and the classifier is a ModernBERT model scoring intent and complexity. It is Apache 2.0, and it ships semantic caching, PII detection and a prompt guard alongside the routing.
The decision it makes is the interesting part. Rather than only choosing between models, it decides whether reasoning is worth enabling for this query — auto reasoning-mode adjustment. The project's published MMLU-Pro figures with a Qwen3 30B model are a 10.2% accuracy improvement alongside a 47.1% latency reduction and 48.5% fewer tokens. Accuracy going up while tokens go down is the tell: on easy queries, reasoning is not merely expensive, it is sometimes actively worse. That is the same argument as adaptive thinking and effort budgets, implemented at the proxy.
What a router actually decides
Compare the three decisions on the same axis and the split is not about accuracy, it is about what question is being asked. A strong/weak threshold asks "is this hard?" and gives you one dial you can explain to a finance team. A learned per-query preference asks "which of these models tends to do best on queries like this?" and gives you breadth at the price of a decision you cannot reconstruct. An effort classifier asks "how much computation does this deserve?" — and on reasoning-model traffic that is where the money is, because the spread between a short answer and a long chain of thought dwarfs the spread between two vendors' per-token prices.
The published savings ranges — commonly quoted between 30% and 85% — are all real and all measured on someone else's traffic. A router's return is a property of your query mix: it pays when a large head of your requests are genuinely easy and a small tail is genuinely hard. Chat assistants look like that. Extraction pipelines do not, because every request is the same difficulty and the right answer is to pick one small model and stop paying a classifier to re-discover that. Measure your own distribution before you believe a percentage; unit economics is the frame, and cost per successfully completed task is the number.
The evaluation you now owe
This is the part that gets skipped, and it is the reason router projects quietly stall six months in. Every other component in your stack fails loudly. A router does not: when it sends a hard query to the small model, you get a 200, a fluent answer, and a slightly worse outcome. There is no error rate to alert on, no latency spike, no log line that says this was the wrong choice.
- Router quality is invisible in ops metrics and visible only against a held-out set with known-good answers. If you do not have one, the router is unfalsifiable and its savings are the only number you will ever see — which is precisely the failure shape that makes it feel like a success.
- Every model release moves the decision boundary. A learned router trained on last quarter's model pair is making calibrated decisions about capabilities that have changed. Routers are perishable in a way that gateways are not, and quality regression detection is the thing that tells you the fruit has turned.
- Log the decision, not just the outcome. Store which model was chosen, by which router version, with what score, on every request. Without that field you cannot answer "did quality drop because of the router or because of the model?", which is the first question anyone will ask.
- Shadow before you switch. Run the router in parallel on a slice of traffic, send everything to the strong model anyway, and compare what the router would have chosen against what the strong model actually produced. The cost of that experiment is a fraction of the savings, and it is the only way to size the regression in advance.
- Budget the hop. A hosted router adds a network round trip before the model call — typically tens of milliseconds, occasionally much worse — and a dependency that can fail. Decide now what happens when the router is down: defaulting to the strong model is right, defaulting to the cheap one is a silent quality incident, and graceful degradation is where that choice belongs.
Routing inside an agent loop breaks the cache
Everything above applies to chat traffic. Agents change the arithmetic, and not in the router's favour.
An agent step re-sends the whole transcript: system prompt, tool definitions, every prior tool result. That prefix is stable and large, which is exactly what prompt caching is for, and cached input reads are typically discounted by around 90% against fresh input. Caches are per-model. Route step 4 to a different model than step 3 and the prefix is cold again — you pay full price for the entire accumulated transcript in order to save a fraction on one step's tokens. In most agent loops the cache discount is the larger number, which means per-step routing can cost you money while every dashboard says it is saving some. See prompt caching and context caching economics for the arithmetic.
There is a correctness argument too, and it is the stronger one. Models differ in how they emit tool calls, how eagerly they call them, how they format structured output, and how they behave when a tool returns an error. Swapping models mid-trajectory means the run's behaviour changes halfway through for reasons unrelated to the task, and reproducing a failure requires knowing which model handled which step. That is a real debugging tax for a saving that the cache has usually already eaten.
The placement that does work is the task boundary rather than the step boundary: classify the incoming task once, pick a model for the whole run, and keep the prefix warm for its duration. That preserves the cache, keeps the trajectory coherent, and still captures most of the spread — because task difficulty, not step difficulty, is what determines whether you needed the frontier model at all. It is also the placement the router pattern describes, and it is why the effort-classifier shape generalises better here than the model-picker shape does.
When to pick which
| If your situation is… | RouteLLM | Not Diamond / Auto Router | vLLM Semantic Router |
|---|---|---|---|
| You want to measure before you commit | Yes — it ships the harness | Hard; the decision is opaque | Yes, on your own stack |
| Already on a gateway, want the easy win | No | This, and no extra vendor | No |
| Self-hosted open weights | Library only | Not the shape | This is the case |
| Reasoning models dominate the bill | Indirect | Indirect | Direct — effort control |
| Agent loops with long transcripts | At task boundary only | At task boundary only | At task boundary only |
| Uniform-difficulty pipeline | Skip it | Skip it | Skip it |
The last row is the one most teams belong in and least want to hear. If every request in a workload is the same shape, the correct router is a decision you make once, in code, for free.
FAQ
Is OpenRouter's Auto Router really the same as Not Diamond?
The Auto Router endpoint uses Not Diamond as its routing engine, so the decision comes from the same place. What differs is packaging: Auto Router arrives inside a gateway you are already calling, with no extra hop, no second contract and no separate key. If you were choosing between them for redundancy, you were choosing between one engine and the same engine.
Do these routers actually cut costs by 30–85%?
Those figures are real measurements on traffic that had a fat head of easy queries and a thin tail of hard ones. Your saving is a property of your own difficulty distribution, not of the router. Measure it on a held-out slice of your own requests before you plan against a number from someone else's workload.
Can I put a router inside an agent loop?
At the task boundary, yes. At the step boundary it fights prompt caching — a model switch makes the accumulated transcript a cache miss, and the discount on cached reads is usually larger than the saving on one step — and it makes tool-calling behaviour change mid-trajectory, which is a debugging cost you pay forever.
What is different about the vLLM Semantic Router?
It runs as an Envoy external processor in your own serving stack rather than as a hosted call, and its primary decision is how much computation a query deserves — including whether to enable reasoning at all — rather than which vendor's model to use. On reasoning-heavy traffic that axis moves far more money than model choice does.
What breaks first when a router goes wrong?
Nothing, visibly, which is the problem. A misroute returns a valid response with a normal status code and slightly worse content. Without a held-out evaluation set and the chosen-model field logged on every request, router regressions are indistinguishable from ordinary model variance.
Further reading
On this wiki:
- Model Routing — the concept, without the products.
- The Router Pattern — where the decision belongs in an agent.
- Prompt Caching — the discount a per-step router throws away.
- Adaptive Thinking & Effort Budgets — the axis the semantic router actually optimises.
- AI Gateways — the layer these either sit inside or in front of.