All four are Apache-2.0 and all four ship PPO and GRPO, so neither the licence nor the algorithm list decides anything. What decides it is whether your environment is a separately scheduled participant in the rollout or a callback inside the generator — because every fix for a slow tool buys throughput by training on stale data. Pick on which staleness knob you get, not on whose speedup number is biggest.
The harness crossed the air gap; the model did not
IBM made self-hosted Bob generally available on 1 October 2026 for on-premises, private-cloud, sovereign-cloud and air-gapped environments, with the shell, parallel tool calling, skills and modes intact. The models supported on customer-managed infrastructure are NVIDIA Nemotron and Poolside Laguna — not the hosted Claude, Gemini and GPT options. The feature list ports; the behaviour has to be re-earned, which makes a sovereignty migration an eval migration wearing infrastructure clothes.
All four ingest OpenTelemetry, so "OTel support" decides nothing — the vocabulary that would make a trace portable is still entirely at Development stability. Pick on who owns the write path and the bulk read path, because production traces are the one asset you cannot re-create, and the licence badge is orthogonal to whether you can get them back.
GTIG reported on 30 September 2026 that exactly 50% of AI-discovered vulnerabilities yield remote code execution against 26% of everything else — but publishes no sample size, and its attribution method selects for the few vendors currently pointing agents at memory-unsafe systems code. The number worth acting on is four sections down: 782 CVEs in agent frameworks and orchestration in eight months, against 97 for frontier models.
A dropped subscription looks exactly like a quiet week
OpenAI shipped plugin automations on all plans on 29 September 2026 against MCP Events — a draft with no SEP number, in a repository whose README calls its contents exploratory. The draft gets the webhook hardening right and makes the two envelopes that report absence optional, so a revoked permission, a lost buffer and a genuinely quiet upstream all reach your agent as the same empty stream.
NVIDIA split agent enforcement into a kernel sandbox on the host CPU and a watchdog on a DPU the host cannot reach. The sandbox is Apache-2.0 on GitHub today; the watchdog has no ship date. The split is not a release accident — the layer far enough away to be tamper-proof is too far away to understand what the agent was trying to do.
A sandbox that permits outbound GET and nothing else reads as a read-only window. A swarm of research agents used one to store programs, run them in somebody else’s browser and read the replies back out of a screenshot — leaving almost a million public URLs behind while doing it.
Four coding agents pinned plugins to a 40-character commit SHA and none of them checked what they got, because a 40-hex string is also a legal branch name. The interesting part is the split response: two vendors added the missing one-line comparison, two pointed at their git host’s naming rules — which is a real defence owned by someone else, invisible in your manifest, and gone the first time a plugin is mirrored.
All four delete the long-lived key in your agent’s environment variable, and the choice between them comes down to where the trust anchor lives and whether humans and machines need one policy plane. None of them answers the question 2026’s agent incidents are actually about: an SVID proves which process is calling, never which user the turn serves or who wrote the instruction now in the context. Buy the floor, then go buy the second thing.
Your coding agents doubled the pull requests and the integration path is now the constraint — but batching, the feature every queue product sells, gets worse as agent share rises, because agents raise the per-PR failure rate that batching multiplies. Only two capabilities change the arithmetic: bisecting a failed batch, and deriving independent lanes from what a change actually touches. Shortlist on those; everything else is configuration.
Bedrock Knowledge Bases vs Vertex AI Search vs Azure AI Search vs Vectara
You are not buying retrieval quality from a managed knowledge base — you are buying the connector that copies SharePoint's permissions along with its files, and the query path that enforces them per user. Azure's Agents SDK search tool still cannot forward that token, and permission lock-in is the layer that actually holds you.
Ollama vs LM Studio vs llama.cpp vs MLX: the tool call is the whole difference
Four local runtimes, the same weights, four different prompts going in and four different answers to whether a tool call comes back parsed. None of that is throughput, and throughput is the only axis anyone compares.
OpenAI vs Gemini vs Perplexity vs Exa: the research API sells you the loop
A search API returns documents and leaves the agent loop in your process. A research API takes the loop, and that is the trade — you stop paying to orchestrate and you stop being able to instrument. The axis nobody tables is what a citation is: three of these four hand back a bibliography the model assembled, and one binds grounding to a field in a schema you defined, with a confidence. Pick on that, not on report quality.
Full duplex deletes the turn — and the turn was your commit point
GPT-Live-1 landed in the API on 10 September and listens while it speaks, which reads as a naturalness upgrade and is actually a schema change. End-of-turn was the event your voice agent used to decide when to call a tool, when to write a log line, when to run a guardrail and when to stop the meter — and a full-duplex model never fires it. The fix is not a better threshold; it is naming your own commit points and pricing a meter that now runs on wall clock instead of speech.
Cursor Projects vs Codex cloud vs Claude Code on the web vs Jules: buy the meter
Four cloud coding agents that look interchangeable on a feature table bill in four different shapes — a usage pool with overage, one allowance shared across every surface you use, a rate limit shared with the rest of your account, and hard task counts per tier — and each shape induces a specific, predictable misuse. Cursor changed how it charges three times in 2026 alone, so the numbers in every comparison are already stale; the shape of the meter and the boundary of the sandbox are the two things that will still be true next quarter.
The Agents API sells you the harness — compaction included
OpenAI opened the Agents API in public beta on 10 September, putting the managed Codex harness — sessions, subagent orchestration, recovery and context compaction — behind one API call, with no fee beyond tokens and containers. The compaction step is the part worth arguing about: it is the transformation that quietly rewrites what your agent is trying to do, and it now runs on a version you cannot pin, diff or roll back. Your eval numbers stop describing a system you control the moment you adopt it.
RouteLLM vs Not Diamond vs vLLM Semantic Router vs OpenRouter Auto
OpenRouter’s Auto Router runs Not Diamond underneath, so four products are three routing decisions. The one that matters for agents is not which model — it is how much computation a query deserves, which is what the vLLM Semantic Router classifies. And a router is a classifier whose errors are silent: it returns a valid, slightly worse answer with a 200, so the savings are the only number you will see unless you keep a held-out set. Inside an agent loop, per-step routing fights prompt caching and usually loses.
All four do hybrid search, graphs and agentic retrieval, so the feature table decides nothing. Two things do: whether the framework runs inside your process or arrives as a second production system with its own database, users and on-call — and where the document-parsing boundary sits, because that is what decides whether your best-quality path is open source, a per-page bill, or an integration you own. Pick the posture; the features converged eighteen months ago.
Wren AI vs DB-GPT vs Vanna vs Dataherald: the generator was never the product
The most-starred open-source text-to-SQL project is read-only — Vanna archived its repo on 29 March 2026 at 23.8k stars — and Dataherald has not taken a commit since July 2024. The two still shipping daily are the two that put a durable, reviewable artefact between the question and the SQL. Frontier models absorbed SQL generation; what they cannot absorb is which of your four definitions of "revenue" this question meant, and that is the layer you own whichever project you pick.
MHS vs SiLA 2 vs OPC UA LADS vs ROS 2: the wire format was never the problem
Lab and factory interoperability has been standardised three times already — SiLA 2 since 2019, OPC UA LADS since January 2024, ROS 2 as robotics middleware — and instruments still ship with vendor SDKs, so a fourth spec is not obviously the answer. What Anthropic's Model Hardware Standard adds is the thing none of the three tried: a device that describes its own limits in language a model can read, and a driver that enforces them whichever model is driving. Useful, and not a safety function — keep those apart.
GraphRAG vs LightRAG vs Graphiti vs Cognee: choose by write pattern
The retrieval quality gap between these four is far smaller than the gap in what an update costs, so the real decision is whether your graph is built once, appended to, or continuously mutated. And all four dedupe entities by string matching, which is the failure nobody's benchmark catches.
Prime Intellect vs HUD vs ART vs OpenAI RFT: you are choosing where the environment lives
Trainers and GPUs are rentable and the base model changes every quarter, so the only durable thing an RL project produces is the environment — the task distribution, the tool surface and the verifier that scores a run. These four platforms disagree about where that artifact lives and who writes the reward, and the one that offered to own the whole pipeline is closing to new users. Pick on portability of the environment and ownership of the verifier; the trainer comparison is the easy part.
E2B vs Daytona vs Modal vs Northflank: the sandbox is idle most of the time
The cold-start number in every pitch deck — 27 ms, sub-90 ms, ~150 ms — describes creating one sandbox at a time. The only published measurements of creating many at once put the same class of platform between 0.67 s and 5.06 s, and two of these four have no published burst figure at all. Meanwhile the sandbox spends most of its life waiting on a model rather than running code, so the axis that actually sets your bill is what the meter does while nothing executes. Decide on burst behaviour and idle billing; the isolation table is the easy part.
agentgateway vs ContextForge vs Obot vs Docker MCP Gateway: whose identity reaches the server
The MCP specification settles the negative — as of revision 2026-07-28 a server MUST NOT pass through the token it received from its client — but leaves RFC 8693 token exchange on the roadmap, so four gateways answer the question four different ways. agentgateway and ContextForge exchange the token; Obot attaches the user’s stored upstream token and is mid-migration between the two; Docker MCP Gateway has no user concept at all, which is honest for a workstation and disqualifying for a fleet. Pick on that axis, and notice that Docker’s isolation story is the best of the four on an axis the others do not compete on.
Temporal vs Restate vs Inngest vs DBOS: where the agent’s transcript lives
All four resume a crashed run from its last completed step, so that is not the decision. An agent’s durable record is a transcript that grows with every turn, not a handful of small step results — and the engines differ on where that growth is stored, what ceiling it hits, and whether your model call is allowed to sit in replayed code. Temporal terminates a workflow at 51,201 events or 50 MB of history; Inngest caps a step output at 4 MB and run state at 32 MB; Restate and DBOS push the growth into storage you operate. Decide on that, then on billing shape, and the feature tables stop mattering.
LiteLLM vs Portkey vs Helicone vs OpenRouter: in the path, or beside it
Two binary questions decide this and no feature list does: is the gateway inside the request path, and who holds the provider credential. Everything a gateway does that changes a request — caching, fallback, rate limiting, key rotation — requires the first, and everything about your blast radius and your bill follows from the second. For agents both answers get multiplied by step count, which is why a choice that is merely fine for a chat app can be structurally wrong for a loop.
Neo4j vs Memgraph vs FalkorDB vs LadybugDB: Picking a Graph Store for Agent Memory
Every performance number published about these four engines was written by one of the vendors, and none of them measures the concurrent-write workload agent memory actually generates. What is checkable — licence, write-path concurrency, and whether your memory framework already ships a driver — points somewhere counter-intuitive.
The 2026-07-28 MCP spec deleted the initialize handshake and the session-id header, so a server can now run behind a plain round-robin load balancer. That operational win is real — but statelessness is a transport property, not a system property. The session did not disappear; its bookkeeping moved onto every request, and the durability that long-running agents actually need came back in through the AWS-contributed Tasks extension as explicit handles. Read the two together before you celebrate a simpler protocol.
A payments company paid a reported $7 billion for the layer that counts AI usage, seven months after buying the layer that invoices it. Routing was never the scarce asset — the scarce asset is one normalised record of what every model call cost and who it was for, and if that record lives in your request path you are paying a percentage on every step your agents take.
ElevenLabs vs Cartesia vs Deepgram vs Rime: buy the tail, not the average
These four advertise time-to-first-audio between 40 and 200 ms, and an independent harness measures their cloud medians at 188 to 313 ms — but the number that breaks a phone call is the spread, not the median, and one vendor's jitter is nearly four times another's. Price moves about 2.5× across the field and predicts neither. The tail is bought with deployment.
Brave vs Exa vs Tavily vs Parallel: the price unit tells you who reads the page
These four price a search between $1 and $16 per thousand, and the spread is not margin — it is how far down the retrieval pipeline each one reads. Price a whole research turn instead of a call and the ordering inverts: the cheapest rate card produces a turn costing three times the dearest one.
gVisor vs Firecracker vs Kata vs WebAssembly: cold start is the operating system
Your sandbox vendor already picked one of these four, and the pick decides whether your agent can run pip install. Rank them by cold start and you get the exact reverse of ranking them by how much Linux the agent gets — because the boot time is the kernel. Answer one question, does the code install things, and the field collapses.
Cloudflare’s Kitesurf makes a browser cheap enough to throw away
The quoted number is 3–7× less CPU and memory than Chromium. The consequence worth planning around is that a fresh browser per task stops being a cost you amortise by reusing sessions — and session reuse is where browser-agent state leaks live. What you trade for it is a compatibility tail that fails silently.
OPA vs Cedar vs OpenFGA vs SpiceDB: who is trusted to supply the facts
All four can express the policy. Only two of them answer without the caller supplying the facts — and when the caller is an agent reading attacker-controlled text, that is the entire security property. The second question is the check budget: an agent makes dozens of authorization calls per task, and filtering a retrieval set makes thousands.
Together vs Fireworks vs Baseten vs Modal: Agents Break Per-Token Pricing
A chat product needs several hundred concurrent users before a dedicated GPU beats per-token pricing; an agent needs about a dozen workers, because it re-sends its whole context every step. That arithmetic — not the price per million tokens — is what should decide which of these four you build on.
Deepgram vs AssemblyAI vs ElevenLabs vs Speechmatics: You Are Buying a Turn Detector
Word error rate is close to settled between the four, and a couple of points of it lands on words your intent classifier ignores. The slice that decides whether a voice agent feels human is end-of-turn detection — several times larger than the transcription latency beneath it, and the one thing the four providers genuinely disagree about.
Arcade vs Composio vs Pipedream Connect vs Nango: Who Holds the User’s Token
Four platforms that stand between your agent and a user’s Gmail or Salesforce. The catalog sizes they advertise are counted in four different units and are the part you will outgrow; the token vault none of them market is the part you would hate to build. The decision you cannot retrofit is whose name is on the consent screen.
AgentCore vs Foundry vs Vertex AI Agent Engine vs Cloudflare Agents: Nobody Is Selling You the Loop
Two of the four bill the agent loop at about nine cents per vCPU-hour and their prices are 3.6% apart; the other two do not charge for it at all. What each is actually selling is a place to keep the conversation — and AWS closing Bedrock Agents Classic to new customers on 30 July 2026 is the clearest evidence yet about which half of a managed runtime you can afford to rent.
Browserbase vs Steel vs Hyperbrowser vs Anchor Browser: You Are Choosing Who Holds the Session
All four speak CDP, so the automation code ports in a day and the SDK comparison decides nothing. The real choice is who holds the logged-in profile, the credentials that recreate it and the exit IP whose reputation you inherit — plus the fact that the latency spread between them is entirely control plane.
Langfuse vs LangSmith vs Phoenix vs Braintrust: The Meter Is the Product
The feature grids converged, so the decision is licence and billing meter — and every meter prices the trace archive that becomes your golden set, regression baseline and fine-tuning corpus. Instrument against OpenTelemetry, dual-write the stream somewhere you own, and the platform becomes a swappable backend.
Auth0 vs Descope vs Stytch vs WorkOS: Agent Auth Is Two Products
Every identity vendor now sells “auth for AI agents”, and the phrase covers two opposite problems: letting an agent into your app, and letting your agent out to someone else’s API. Pick on direction first — and notice that a token vault holding a user’s full grant has relocated the credential rather than shrunk it.
E2B vs Daytona vs Modal vs Cloudflare Sandboxes: Pick on the Billing Shape
A sandbox serving a twenty-step agent spends about six sevenths of its life idle, waiting for a model to think. So cold-start milliseconds and per-vCPU-hour rates — the two numbers every comparison leads with — are the two that matter least. What decides your bill is whether idle is billed; what decides your blast radius is the egress default.
Cohere vs Voyage vs Jina vs Qwen3: The Retrieval Model You Can Actually Un-Choose
A reranker touches no index and holds no state, so swapping one is an afternoon — which finally makes chasing the leaderboard rational, except the leaderboard measures the axis where these four differ least. What differs by more than an order of magnitude is the billing unit and the licence, and both bite hardest at agent scale.
Outlines vs XGrammar vs llguidance vs Instructor: Valid JSON Was Never the Hard Part
Three of these four constrain the sampler so invalid output cannot be produced, and the choice between them collapses to one question: do your schemas repeat? The fourth does something categorically different, and it is the only one that can enforce the rules that actually break agents — because a grammar guarantees the enum is one of five values and says nothing about which.
vLLM vs SGLang vs TensorRT-LLM vs llama.cpp: Throughput Is the Wrong Benchmark for Agents
Every comparison of these four opens with tokens per second on a fixed batch — the one number that transfers worst to agent traffic, where the same prompt comes back twenty times with a few hundred tokens appended. What separates them is what the KV cache is keyed on, whether constrained decoding survives a full batch, and how much of your quarter the build step eats.
The 28 July specification retires the initialize handshake and the Mcp-Session-Id header, and every write-up so far has framed that as plumbing. It is not. Dropping the held-open connection forced Sampling, Roots and Logging onto a twelve-month deprecation clock — and those were the features that made an MCP client a peer rather than a caller. The protocol just settled what it is.
LiteLLM vs Portkey vs Cloudflare AI Gateway vs Kong AI Gateway: Four Bets on What Sits Between Your Agent and the Model
Every AI gateway sells the same headline feature: automatic failover to a second provider. That feature is not an availability win — it is an untested deploy that fires only during an incident, onto a model your evals never covered. Choose instead on who operates the hop, because that is the decision you cannot reverse cheaply.
Exa vs Tavily vs Brave Search vs Firecrawl: Four Bets on How an Agent Should Search the Web
List prices for agent search APIs cluster tightly around $5–8 per thousand queries, which makes the sticker the least interesting number in the comparison. What actually differs by an order of magnitude is how many tokens each one dumps into your context per result — and in an agent loop that re-sends its transcript every step, that is the bill.
Kimi K3 Is Open Weights. That Is Not the Same as Cheap, Local, or Unrestricted
Moonshot released 2.8 trillion parameters as a free download on 27 July — and priced its own API above the model it replaced, while no single GPU on the market can hold the weights. Open weights buy agent builders exactly one thing that closed APIs cannot, and it is not cost.
Temporal vs Inngest vs Restate vs Cloudflare Workflows: Four Bets on Keeping Your Agent Alive for 30 Minutes
Naive agent loops die on minute 29 of a 30-minute job. Durable-execution engines journal every step so the next process can pick up exactly where the previous one died — and 2026 was the year hyperscalers shipped their own. Four engines now compete on the same primitive, with very different architectures and bills.
Mem0 vs Letta vs Zep vs Cognee: Four Bets on What "Agent Memory" Actually Means
A 128K-token context window degrades past the first thousand tokens and vanishes the moment the session ends. The agent-memory infrastructure market crossed $6 billion in 2026 because "throw it all in the context" stopped being a strategy — and four frameworks now bet differently on what memory should rank, store, and forget.
MCP at 97 Million Downloads: How the Model Context Protocol Won — and What's Still Broken at Scale
Two years from Anthropic's launch, MCP isn't a debate — it's a dependency. Every frontier vendor, every major IDE, and one Pinterest team saving 7,000 engineering hours a month all ship against it. The interesting question is no longer *should you use MCP* — it's what fails at this scale and how the 2026 roadmap plans to fix it.
pgvector vs Pinecone vs Weaviate vs Qdrant: Where the Index Sits Decides Everything
Four vector stores, four nearly identical feature lists — ANN, filters, hybrid search, all of it. The thing that actually decides which one survives the agentic-RAG stack at scale is invisible there: where the index sits relative to your primary data.
LangSmith vs Braintrust vs Helicone vs Arize Phoenix: Four Loops the Eval/Observability Stack Was Built to Close
All four ship traces, datasets, and evaluators — the feature lists nearly match. What separates them is which feedback loop they were built to close: the dev loop, CI, the production gateway, or model-monitoring drift.
E2B vs Modal vs Daytona vs Anthropic Code Execution: Four Owners of the Agent Sandbox
Four runtimes give an agent a place to actually execute Python and bash safely — and the marketing pages all promise the same thing. The thing that decides which one survives production is who owns the sandbox lifecycle.