AI Blog

vLLM vs SGLang vs TensorRT-LLM vs Dynamo: the second replica is where your cache went

Published head-to-heads on these engines range from a 29% edge to a 6.4x one, because none of them state the parameter that decides an agent workload: how much of each request is a prefix you already paid for. Three of these four are engines and one is a router — and the moment you add a second replica, your cache hit rate stops being an engine property at all.

By Agentic AI Wiki 12 min read

Published head-to-heads between these projects put the winner's margin anywhere from 29% to 6.4×, on the same hardware class, for the same models. Both numbers are probably honest, and the reason they disagree by two orders of magnitude is that neither states the one parameter that decides an agent workload: how much of each request is a prefix you already paid to compute. Pick on that, and you discover the uncomfortable thing about this comparison — three of these four are engines, one is a router, and the moment you run a second replica your cache hit rate stops being an engine property at all.

At a glance

Four projects that get compared as peers, which three of them are.

ProjectWhat it isLatest releaseWhat you are actually choosing
vLLM Inference engine v0.31.0, 5 October 2026 (roughly every two weeks) Breadth: hardware, model coverage, and the default everyone else integrates against
SGLang Inference engine v0.5.21, 2 October 2026 (roughly every two weeks) A token-level radix tree over the prefix cache, now on a Rust core by default
TensorRT-LLM Compiled inference engine, NVIDIA only 1.3.0rc line, still rc as of late September 2026 Peak per-GPU performance, paid for in setup time and vendor lock
Dynamo Router and serving platform over the engines v1.5.0, its 18th feature release Which worker a request lands on — i.e. whether the prefix cache is hit at all
Prefix cache hit rate by routing posture Horizontal bars comparing reported prefix cache hit rates: a single worker holding the session reaches 85 to 97 percent, cache-aware routing across four replicas reaches 89 percent in one third-party test, and spreading requests evenly across four replicas without cache-aware routing lands near the one-over-N approximation of 25 percent. Reported prefix cache hit rate (%) Session pinned to one worker vendor-reported, coding-agent traffic 85–97 Cache-aware routing, 4 replicas third-party test, ~50k-token inputs 89 Even spread, 4 replicas 1/N approximation, not a measurement ~25 0 25 50 75 100 Same engine, same model, same traffic in all three rows. The variable is routing.
Same engine, same model, same traffic in all three rows. The variable is routing.

Hugging Face's TGI is the fourth engine that used to be in this table, and its absence is the most informative data point available. It went into maintenance mode in December 2025, the repository was archived read-only on 21 March 2026, and its own README now points readers at vLLM and SGLang. Nothing about its throughput got worse. The axis moved.

One of these four is not an engine

Where the router sits relative to the engines Agent traffic enters a router that holds a global index of which key-value cache blocks live on which worker. The router dispatches to workers, each running one engine — vLLM, SGLang or TensorRT-LLM — and each holding its own local prefix cache. Without the router, requests are spread across workers and the prefix cache on any one worker is missed. Agent traffic system prompt + tools + history, every turn Router (Dynamo) global index: which KV blocks on which worker No router turn 2 lands ~1/N on the right worker Worker 1 — vLLM block-level hashed prefix cache v0.31.0 · 2026-10-05 Worker 2 — SGLang token-level radix tree, Rust core v0.5.21 · 2026-10-02 Worker 3 — TensorRT-LLM compiled engine, KV reuse 1.3.0rc · NVIDIA only Each worker’s prefix cache is local. Reuse is whatever your routing delivers to it. Dynamo v1.5.0 pins SGLang v0.5.18, TensorRT-LLM v1.3.0rc25, vLLM v0.28.0 — the router sets the engine calendar.
Dynamo does not serve tokens. It decides which worker's cache your request meets.

The clearest evidence that Dynamo belongs on a different row is its own release notes: v1.5.0 pins SGLang v0.5.18, TensorRT-LLM v1.3.0rc25 and vLLM v0.28.0 as backends. It runs the other three. Comparing it to them on tokens per second is a category error, and comparing the other three to each other without saying what sits above them is the more common and more expensive one.

That pin list carries a second, quieter cost worth pricing before you adopt the platform: as of early October 2026 upstream vLLM is at v0.31.0 and SGLang at v0.5.21, several releases past what v1.5.0 pins. Taking the router means taking the router's engine calendar, which trails upstream by several releases on a two-week cadence. If you are on a self-hosted stack precisely because you need a model or kernel that landed last fortnight, that gap is the trade you are making, and it is not mentioned in any benchmark.

Why the published comparisons disagree by 6×

Read a handful of 2026 head-to-heads and the spread is startling. One reports TTFT at the median dropping from 310 ms to 195 ms at 80% shared prefix and 50 concurrent requests. Another reports up to 6.4× higher throughput on prefix-heavy traffic, while noting in passing that a 29% figure from the same family of tests applies to the workload that least favours the winner. A third finds one engine's multi-turn throughput sagging under cache pressure where the other holds steady.

These are not contradictory results; they are three different workloads wearing the same label. The missing axis is prefix reuse — the share of each request's tokens that the server has already computed for some earlier request. Agent traffic sits at the extreme end of it: a system prompt, a tool catalogue and a growing conversation history are re-sent on every turn, so the marginal request is mostly a prefix. One write-up puts agentic workloads with identical tool definitions at 75–95% reuse; that specific number is untraceable to a measurement, but the shape is not in dispute, and NVIDIA's own coding-agent example reports 85–97% cache hits on subsequent calls that reach the same worker.

So the benchmark that would decide your choice is one nobody publishes, because it is yours: what fraction of your tokens are prefix, and what is your current hit rate? Both are already in your server metrics. An engine comparison run at 0% reuse measures a regime your agent never operates in, and an engine comparison run at 90% reuse mostly measures the cache implementation — which is why the same two projects can swap places depending on who ran the test.

The second replica is where the cache goes

Here is the part that reorders the whole decision. A prefix cache is local to a worker process. With one replica, every turn of a session necessarily meets the cache that the previous turn populated, and you get the headline hit rates. With N replicas behind an ordinary load balancer, turn two lands on the worker that holds turn one's prefix with probability roughly 1/N — NVIDIA's own framing of the baseline, and an approximation for random placement rather than a measurement, but directionally exactly right.

Which means scaling out actively destroys the property you chose your engine for. Four replicas, round-robin, and a workload that was 90% prefix is now recomputing most of that prefix most of the time. Your throughput chart goes up because you added GPUs; your cost per task goes up too, and nothing in the engine comparison warned you.

Cache-aware routing is the fix, and it is a layer, not a setting. Dynamo's router maintains a global index of which KV blocks live on which worker, scores each candidate worker for overlap against the incoming request, and picks the one that minimises the combined cost of a cache miss and the worker's current decode load. One third-party test — Baseten, on Qwen3 Coder with roughly 50k-token inputs — reports an 89% hit rate across four replicas with a 50% reduction in time-to-first-token and 34% in time-per-output-token. Vendor-adjacent numbers, best-case workload, but the mechanism is sound and the baseline it beats is arithmetic rather than opinion.

Three things about that layer you should know before you depend on it:

  • Stickiness and load balance are in tension, by design. The overlap credit is a tunable, with a documented default that NVIDIA calls a reasonable starting point rather than an optimum; push it toward cache affinity and you concentrate traffic on cache-rich workers. There is no setting that gives you both.
  • Tie-breaking is random. When two workers score identically the router picks among them at random, and there is an open report of exactly that case in a multi-turn chatbot with two decode workers holding matching blocks. In a two-replica deployment, "identical scores" is not a rare state.
  • Multi-tier caches get partial credit. Overlap found in CPU-offloaded cache counts at a fraction of GPU-resident overlap, and disk or NVMe at a smaller fraction still. Your effective hit rate is not one number.

What actually differs at the engine layer

Where each project leans hardest A matrix of four projects against four axes. Rows are vLLM, SGLang, TensorRT-LLM and Dynamo. Columns are prefix cache mechanism, hardware and ecosystem breadth, setup cost, and cross-replica cache routing. Dynamo is the only row strong on cross-replica routing and the only one that is not itself an engine. Prefix reuse, breadth, setup cost, cross-replica routing prefix cache breadth cheap setup cross-replica vLLM block hash strong strong not its job SGLang radix tree medium medium not its job TensorRT-LLM KV reuse NVIDIA only weak not its job Dynamo delegated 3 backends weak strong leads on this axis competent not where it competes Only one row is strong in the right-hand column, and it is the row that is not an engine.
Where each project leans hardest — and which column only one of them competes in.

With routing accounted for, the engine differences get smaller and more legible. Both of the open engines cache prefixes; they differ in how they index them. vLLM hashes fixed-size blocks, which is simple, cheap to look up, and aligned to its paged allocator. SGLang keeps a token-level radix tree, which shares partial prefixes more finely — and it has been investing there specifically, moving the prefix cache onto a Rust core by default in v0.5.21 after unifying the radix tree in the two releases before it. On traffic with many near-identical-but-not-identical prefixes, finer sharing is worth real money; on traffic where prefixes are identical or completely different, the mechanisms converge.

Breadth runs the other way. vLLM is where new hardware, new model architectures and third-party integrations land first, and it is the engine most other tools assume. SGLang's release notes read like a project optimising hard in one direction, which is a compliment and a risk assessment at once. TensorRT-LLM is the peak-performance option on NVIDIA silicon and asks for a compile step, a week or two of setup, and acceptance of single-vendor hardware; its 1.3 line was still shipping under release-candidate tags as of late September 2026, which is worth knowing if your change-management process wants to pin a stable version number.

When to pick which

SituationPickWhy
One or two GPUs, one model, agent traffic vLLM At single-replica scale the cache is hit by default; take the breadth and the biweekly releases, and spend your attention elsewhere
Prefix-heavy traffic with many partial-overlap variants SGLang Token-level radix sharing is the mechanism that matches that workload, and it is where the project is investing
Fixed model, NVIDIA fleet, latency is the product TensorRT-LLM Peak per-GPU numbers are real; budget the setup and accept the hardware lock
Three or more replicas serving sessions or agents Dynamo over any of the above Beyond two replicas the dominant term is routing, not the engine — and no engine choice recovers a 1/N hit rate
Still on TGI Migrate Archived read-only since March 2026; its own README sends you to vLLM or SGLang

FAQ

Is Dynamo a replacement for vLLM or SGLang?

No. It runs them as backends — v1.5.0 pins specific versions of SGLang, TensorRT-LLM and vLLM — and its job is routing, disaggregation and fleet-level concerns. If you see it benchmarked head-to-head against an engine, the comparison is measuring a stack, not a peer.

Which engine is fastest for agent workloads?

Not answerable without your prefix reuse rate, which is why the published answers range from 29% to 6.4×. Measure the share of your tokens that are prefix and your current cache hit rate first; below a couple of replicas the engine choice dominates, above that routing does.

Do I need cache-aware routing with only two replicas?

Probably yes for sessions, and the caveat matters: at N=2 an even spread costs you about half your potential hit rate, but it is also the configuration where identical overlap scores are most common, and the router breaks ties at random. Verify your actual hit rate after enabling it rather than assuming it.

Is TGI dead?

Effectively. Maintenance mode from December 2025, repository archived read-only on 21 March 2026, and Hugging Face's own documentation now recommends vLLM and SGLang. Existing deployments keep running; new work should not start there.

Why does TensorRT-LLM not have a stable 1.3.0?

Its 1.3 line was still shipping under release-candidate tags as of late September 2026, and Dynamo pins an rc build as its backend. This is normal for the project's cadence but it is a genuine friction if your release process requires pinning a non-rc version.

What is the one number I should measure before choosing?

Prefix cache hit rate, segmented by replica count, on your own traffic. It is already exported by every engine here, it is the parameter all the public benchmarks leave unstated, and a low value tells you to fix routing before you change engines.

Further reading

On this wiki:

Project sources: