Streamable HTTP: the current MCP transport

16 min read

C4
Deep Dive · MCP

Streamable HTTP is still the MCP transport — but the 2026-07-28 revision took the session header, the standalone GET stream, and resumability back out of it.

Streamable HTTP survived the 2026-07-28 revision; most of what a 2025-11-25 tutorial says about it did not. Mcp-Session-Id is gone, GET on the MCP endpoint is gone, and SSE streams are no longer resumable — what replaces them is three required headers on every POST, per-request negotiation inside params._meta, and one explicitly-requested subscriptions/listen stream for anything the server wants to push. The distributed-systems bill that made the old transport interesting was not paid down; it was deleted. The superseded shape stays on this page in the past tense, because a lot of running servers still speak it.

STEP 1

Deprecation timeline: HTTP+SSE gave way to Streamable HTTP, then 2026-07-28 gutted it.

The MCP transport chapter has been rewritten three times in two years, and confusion about which revision a tutorial is describing accounts for most of the "why doesn't my client talk to my server" tickets that reach the reference SDKs. The 2024-11-05 revision defined an HTTP+SSE transport with two endpoints — clients POSTed JSON-RPC requests to /messages/ and held an open SSE stream at /sse to receive responses and server-initiated notifications. The 2025-11-25 revision deprecated that shape and defined Streamable HTTP: a single MCP endpoint that accepted both POST and GET, that could return a plain JSON response or upgrade to text/event-stream, and that carried sessions and resumability in headers rather than in the URL structure. The 2026-07-28 revision kept the endpoint and deleted most of the rest — no initialize handshake, no session header, no GET, no Last-Event-ID. The 2026-07-28 revision essay has the full removal inventory and the migration path; this essay covers what the transport looks like on the wire now, and what the superseded shape looked like so you can recognise it in a codebase you inherit. The MCP architecture essay sketches the participant model and the JSON-RPC layer this transport carries.

This page is an exhibit in its own argument. It shipped in July 2026 opening with a warning that any tutorial describing a POST-then-SSE two-endpoint dance predated the current spec — and within weeks the current spec had moved again, which made the warning true of the page carrying it. The lesson is not that MCP is unstable so much as that transport content has a short half-life and should be read with a revision date in hand. Check the date on anything you are copying from, including this.

The first migration was fast: Bloomberry's 2026 survey of 1,412 production servers found 93% already on Streamable HTTP within six months of the 2025-11-25 bump. The second is slower, because it is breaking rather than additive. fastmcp for Python, @modelcontextprotocol/sdk for TypeScript, the official Go SDK and rmcp for Rust all implement the current revision; mcp-go — the library most Go MCP servers were actually built on — still implements 2025-11-25, which is the practical reason the superseded shape below is documented rather than deleted. Reading rule of thumb: a post that describes an initialize handshake and a session header predates the current spec, and a post that describes a two-endpoint POST-then-SSE dance predates the one before that.

The wire format was never the interesting story. What was worth the essay is that Streamable HTTP took the transport from something you could hold in your head as "HTTP with server-push" to a system with sticky routing, session storage and a replay contract — and then 2026-07-28 took most of that back out. The rest of this essay walks that arc: what the operational profile was, what survives of it, and which parts you now have to build yourself because the transport stopped doing them for you.

STEP 2

The single endpoint: POST to send, SSE to stream, GET to 405.

A Streamable HTTP MCP server exposes exactly one path — the spec calls it "the MCP endpoint" and lets the server pick the URL, but by convention it is /mcp. Clients POST JSON-RPC requests to that endpoint and inspect the response's Content-Type. If the header is application/json, the body is the JSON-RPC response and the exchange is over. If the header is text/event-stream, the server has chosen to stream — the same response is delivered as one or more SSE data: frames, with any progress or logging notifications for that request interleaved, and the stream closes when the server has nothing more to say. The client cannot tell in advance which mode the server will pick and the spec is explicit that both are valid; well-behaved clients handle both. That much is unchanged.

What changed is that the POST now carries its own introductions. There is no handshake to establish protocol version, capabilities or client identity once per connection, so every request restates them — in headers for anything an intermediary needs to see, and in params._meta for everything else.

POST /mcp HTTP/1.1
Host: mcp.example.com
Content-Type: application/json
Accept: application/json, text/event-stream
MCP-Protocol-Version: 2026-07-28
Mcp-Method: resources/read
Mcp-Name: file:///repo/README.md

{"jsonrpc":"2.0","id":1,"method":"resources/read",
 "params":{"uri":"file:///repo/README.md",
           "_meta":{"io.modelcontextprotocol/protocolVersion":"2026-07-28",
                    "io.modelcontextprotocol/clientCapabilities":{"elicitation":{}},
                    "io.modelcontextprotocol/clientInfo":{"name":"acme-client","version":"2.0.0"}}}}

HTTP/1.1 200 OK
Content-Type: application/json

{"jsonrpc":"2.0","id":1,"result":
 {"contents":[{"uri":"file:///repo/README.md","mimeType":"text/markdown","text":"# acme"}],
  "resultType":"complete","ttlMs":60000,"cacheScope":"private"}}

Three headers are load-bearing on that POST, and none of them is optional the way MCP-Protocol-Version once effectively was. MCP-Protocol-Version (all caps, as the spec spells it) MUST be present and MUST match the io.modelcontextprotocol/protocolVersion value inside params._meta; a disagreement between the two copies is 400 plus HeaderMismatch (-32020). Mcp-Method MUST be present on every request and carries the same string as the JSON-RPC method field. Mcp-Name MUST be present on tools/call, resources/read and prompts/get, carrying params.name or params.uri; a value that is not ASCII MUST use the spec's Base64 sentinel format rather than being stuffed into a header raw.

It is worth saying out loud why those last two are headers at all, because it is the design decision that pays for their verbosity: a proxy, gateway or WAF can route, meter and authorize a request without parsing the JSON body. Per-tool rate limits, method-level blocking and name-based sharding become ingress configuration instead of application code, which is exactly the layer where most organisations already own that policy. The cost is that a request missing either header is not a well-formed request, so a hand-rolled client that only sets Content-Type will fail against a conformant server in a way that looks like a routing bug.

Inside the body, params._meta carries what the handshake used to establish once. io.modelcontextprotocol/protocolVersion and io.modelcontextprotocol/clientCapabilities are REQUIRED, io.modelcontextprotocol/clientInfo SHOULD be sent, and io.modelcontextprotocol/logLevel is optional — it is what survives of logging/setLevel, which used to set a level for a whole session and is gone. Omit a required field and the server answers -32602 with HTTP 400. Use a feature the request did not declare in clientCapabilities and the server answers MissingRequiredClientCapabilityError (-32021), also 400. A few hundred extra bytes per request buys the property that any replica can serve any request with no prior context.

The GET half of the endpoint is gone. A GET — or a DELETE, which used to terminate a session — against a modern-only server SHOULD be answered with 405 Method Not Allowed. Server-initiated push has not disappeared, but the client now has to ask for it by name: subscriptions/listen is an ordinary POST whose response stream is held open for as long as the client wants notifications, with a params.notifications filter that says exactly which ones. The filter has four fields — toolsListChanged, promptsListChanged and resourcesListChanged as booleans, and resourceSubscriptions as a string[] of URIs. That array is where resources/subscribe and resources/unsubscribe went: the subscription set is no longer server-side state you mutate with two RPCs, it is an argument of the call that opens the stream. Building an idiomatic server, in the sense the building MCP servers in practice essay uses that phrase, mostly means letting the SDK manage this stream for you; hand-writing it is rarely necessary and easy to get subtly wrong.

POST /mcp HTTP/1.1
Host: mcp.example.com
Content-Type: application/json
Accept: text/event-stream
MCP-Protocol-Version: 2026-07-28
Mcp-Method: subscriptions/listen

{"jsonrpc":"2.0","id":42,"method":"subscriptions/listen",
 "params":{"notifications":{"toolsListChanged":true,
                            "promptsListChanged":false,
                            "resourcesListChanged":false,
                            "resourceSubscriptions":["file:///repo/README.md"]},
           "_meta":{"io.modelcontextprotocol/protocolVersion":"2026-07-28",
                    "io.modelcontextprotocol/clientCapabilities":{}}}}

HTTP/1.1 200 OK
Content-Type: text/event-stream

data: {"jsonrpc":"2.0","method":"notifications/subscriptions/acknowledged","params":{"_meta":{"io.modelcontextprotocol/subscriptionId":"sub-7c1e"}}}

data: {"jsonrpc":"2.0","method":"notifications/tools/list_changed","params":{"_meta":{"io.modelcontextprotocol/subscriptionId":"sub-7c1e"}}}

Three rules on that stream are easy to get wrong. The server MUST send notifications/subscriptions/acknowledged as the first message and MUST NOT send any notification before it, so a client that starts dispatching on the first frame it sees is reading a stream it has not confirmed is live, and a server that pushes eagerly is non-conformant. Every notification on the stream carries the _meta key io.modelcontextprotocol/subscriptionId, which is how a client holding more than one listen stream attributes what arrives on each. And request-scoped notifications do not travel here at all: notifications/progress and notifications/message stay on the response stream of the request that caused them, which is the only stream that knows whether that request is still running. Note also what is absent from those frames — no id: lines, because SSE event IDs went away with resumability.

STEP 3

Sessions: the session header, and the sticky-routing bill 2026-07-28 cancelled.

Under 2025-11-25 the session header was where the transport's distributed-systems bill first came due. The server generated the id in response to the first initialize POST, returned it as a response header — that revision spells it MCP-Session-Id — and expected to see it on every request the same client made thereafter. What the id addressed depended on the server: any protocol-level state that persisted across requests — subscription lists behind notifications/tools/list_changed, elicitation callbacks awaiting a user reply, per-session sampling budgets, retained SSE event buffers used for resumability — was looked up by session id. A stateless server could decline to issue one, but opting out meant losing resumability and the server-initiated flows, since those inherently spanned requests. The 2025-11-25 spec was explicit that session ids MUST be cryptographically secure, opaque to the client, and encode no user data.

2026-07-28 removed the header outright, and with it the concept. A modern-only server SHOULD ignore an incoming session header rather than erroring on it, and MUST neither mint nor echo one of its own. (A small note for anyone grepping: the 2026-07-28 backward-compatibility section spells it Mcp-Session-Id, where 2025-11-25 spells it MCP-Session-Id. HTTP header names are case-insensitive, so this is a documentation artefact rather than a wire difference — but quote whichever casing the revision you are citing uses.) There is no per-connection state now and no replacement for it. The spec's wording is blunt: state that spans requests "MUST be referenced by an explicit identifier the client passes on each request", and "an open connection, such as a STDIO process, is not a conversation or session." If your server kept a cursor, a workspace selection or a half-built transaction between calls, that thing now needs a name, a lifetime, and an id travelling in the request body.

The rest of this step is maintenance reading — it applies to the 2025-11-25 servers still in production, not to anything you are writing today. The moment such a server has more than one replica, the session id becomes a routing problem: if replica A issued the session and holds the state it addresses, replica B cannot serve a request against that session unless the state is shared, which means either the load balancer pins the session to a replica or a shared store (Redis, DynamoDB, an equivalent) holds it for every replica to read. Load balancers that hash on the client IP address are the pattern most teams reach for first and the one that fails most predictably: mobile clients change IPs, corporate NAT collapses many users onto a few addresses, and cloud-hosted clients rotate egress across a pool. Hashing on the session header directly — supported by every ingress controller worth using — is the correct primitive when a pin is what you want. If you are writing a new server, the correct hash key is no key at all.

STEP 4

Resumability: Last-Event-ID, the SSE replay contract, and its removal.

Resumability was the feature that most obviously separated 2025-11-25's Streamable HTTP from a naive "HTTP with events" transport. When the server chose to stream, each SSE frame carried an id: field — an opaque, server-assigned identifier the client was expected to remember. If the connection dropped before the stream completed, the client reconnected by opening a new GET on the MCP endpoint with the standard SSE Last-Event-ID header set to the id of the last frame it had successfully received; the server was then obligated to replay every event issued after that id, in order, before continuing with fresh ones. The client experience was a seamless stream across TCP resets, load-balancer timeouts, and browser tab suspends.

# Superseded: the 2025-11-25 resume. Kept here to be recognisable, not copied.
GET /mcp HTTP/1.1
Host: mcp.example.com
Accept: text/event-stream
MCP-Protocol-Version: 2025-11-25
MCP-Session-Id: 3f9a2c81-0b6d-4e2a-9b71-8d5e2c0a4f11
Last-Event-ID: evt-00047

HTTP/1.1 200 OK
Content-Type: text/event-stream

id: evt-00048
data: {"jsonrpc":"2.0","method":"notifications/progress","params":{"progressToken":"pt-9","progress":0.7}}

# The same request against a 2026-07-28 server:
HTTP/1.1 405 Method Not Allowed
Allow: POST

2026-07-28 removed SSE event IDs and Last-Event-ID together, which means streams are simply not resumable. A modern-only server ignores an inbound Last-Event-ID; there is no retained buffer to address and nothing to replay from. When a stream breaks, the client MUST re-issue the work as a new request with a new request ID — not a reconnect, a new request — and for a subscriptions/listen stream it MUST re-send subscriptions/listen, filter and all. The consequence lands on tool authors rather than transport authors: a re-issued request is a genuinely new request, so anything with a side effect needs to be idempotent or to carry a client-supplied dedupe key in its arguments. That was already true and easy to skip while replay papered over it. It is not skippable now.

The replay budget — how long a server retained post-id events so a reconnect could succeed — was the one number the 2025-11-25 contract needed and never specified, and its disappearance is the clearest thing the removal bought. Nobody sizes a retention window any more; nobody exhausts memory on a session that stays open all day while the client hardly reads from it; and the classic bug where event ids came from an in-process counter that reset on deploy — making resumability silently useless, because the client reconnected with an id the new process did not recognise, and the missing events were simply lost — can no longer happen. What you get instead is a cheaper server and a client that has to be explicit about what it is willing to redo. The teams who lose most in that trade are the ones streaming very long tool calls, and that case is exactly what the io.modelcontextprotocol/tasks extension exists to carry: durable work now gets an identifier you poll, rather than a stream you hope survives.

STEP 5

Horizontal scaling: what broke when sessions were stateful.

Once resumability was real, horizontal scaling stopped being a matter of adding replicas and became a design problem. The load balancer had to route the same session's traffic to the same replica (sticky routing) or every replica had to be able to serve any session (shared state); the SSE reconnect carrying Last-Event-ID had to land on a process that either held the referenced event or could read it from wherever events were persisted; the session TTL had to be coordinated across replicas so that one did not garbage-collect a session another was still serving. Every one of these broke under naive scaling patterns, and the failure symptoms — reconnects that got empty streams, elicitation replies that never resolved, sampling requests that hung — looked like protocol bugs to a client author but were really deployment bugs.

Two shapes dominated the field. The sticky-routing shape kept session state in memory on the replica that owned the session, used the session header as the load-balancer hash key, and accepted that a replica loss meant every session on that replica had to reinitialize; it was simple, had low tail latency, and scaled linearly until a hot session became a replica-level hotspot. The shared-state shape stored session data in Redis or an equivalent, let any replica serve any request, and paid a round trip on every request in exchange for the freedom to move traffic around; replicas became interchangeable and rolling deploys stopped dropping sessions. Teams picked sticky until an incident forced them to shared, and that migration was a well-known milestone in the life of a Streamable HTTP server. 2026-07-28 deletes the milestone. A request is self-describing, so plain round-robin works, a deploy is an ordinary rolling restart, there is no session store to keep available, and the gateway can shard on Mcp-Method or Mcp-Name without opening the body at all.

What remains is smaller but not nothing, and the phrase "remote MCP server" is still doing more work than most tutorials admit. A subscriptions/listen stream still pins one long-lived connection to one replica for its lifetime, so connection counts, idle timeouts and the blast radius of a restart are still capacity planning; the difference is what happens after a restart, which is nothing much, because there is no state to migrate and nothing to replay — the client re-sends subscriptions/listen and is whole again. Anything genuinely stateful now lives behind an explicit identifier you designed, or in the io.modelcontextprotocol/tasks extension, where its durability is your problem in a place you can see it rather than an implicit property of a TCP connection. A local stdio MCP server is a subprocess your host launches; a remote Streamable HTTP server under 2025-11-25 was a distributed system with sessions, sticky routing and replay buffers; under 2026-07-28 it is much closer to an ordinary HTTP API whose remaining hard parts are authorization, rate limiting, multi-tenancy and those listen streams. That is a genuine simplification, and it is also a reason to re-read your own deployment: a lot of MCP infrastructure was built to solve problems the protocol no longer has.

STEP 6

Local HTTP servers: Origin, DNS rebinding, and the localhost trap.

Streamable HTTP is not only a remote-server concern, and this is the one step of the six the revision barely touched — Origin checking and socket binding are transport-level defenses that neither revision changes. A common pattern for local MCP servers is to bind an HTTP listener to 127.0.0.1 instead of speaking stdio, because it is easier to reuse in browser-based clients or to develop against with curl. This pattern comes with a specific attack — DNS rebinding — that MCP's security best-practices appendix has called out by name since the 2025-11-25 revision, because it has been demonstrated in the wild against locally running dev tools. The mechanic is simple: an attacker gets the user's browser to load a page from an origin the attacker controls, that origin's DNS resolves to a public IP just long enough to serve the page, and then rebinds to 127.0.0.1. From the browser's point of view the origin has not changed, same-origin checks succeed, and JavaScript in the attacker's page can now issue same-origin requests against the user's local MCP server, invoking any tool it exposes.

Two defenses stack, and the spec expects both. The first is to check the Origin header on every request and reject anything whose origin is not on an explicit allowlist — typically the host running the local UI. The second is to bind the listener to 127.0.0.1 or ::1 explicitly, not to 0.0.0.0 or an empty host string, so that even if the Origin check is misconfigured the socket is unreachable from any interface but loopback. Servers that skip either defense have historically been exploitable within the first week of shipping; the pattern is not theoretical.

// Express-shaped Origin allowlist middleware for a local Streamable HTTP server.
const ALLOWED_ORIGINS = new Set([
  'http://localhost:5173',
  'http://127.0.0.1:5173',
  'https://claude.ai',
]);

app.use('/mcp', (req, res, next) => {
  const origin = req.headers.origin;
  if (!origin || !ALLOWED_ORIGINS.has(origin)) {
    return res.status(403).json({ error: 'origin_not_allowed' });
  }
  next();
});

app.listen(3945, '127.0.0.1'); // bind loopback explicitly, never 0.0.0.0

Two footnotes on the snippet. First, an empty or missing Origin header is not a safe default to accept; browsers include the header on cross-origin requests, and the absence of it is either a non-browser client (which should be handled through auth, not through Origin) or a browser making a request the server shouldn't be serving in the first place. Second, the allowlist is not a stand-in for authentication — a local server that binds loopback and checks Origin still needs the same OAuth 2.1 story that a remote server does if the tools it exposes have any consequence. Origin checks defeat DNS rebinding; they do not defeat a compromised local process, a malicious host config, or the token-passthrough anti-pattern that the security best-practices appendix calls out separately. The transport-level defenses are one layer; the auth-level defenses are another; a production local server ships both.

Reading the six steps together, the Streamable HTTP transport is still two documents in one, and 2026-07-28 made the first shorter and the second much shorter. On the wire it is a compact HTTP profile — one endpoint, one verb, three required headers, a _meta block, and an optional SSE upgrade — that any HTTP server can implement in an afternoon. In deployment it is no longer the point at which an MCP server becomes a distributed system: sessions, sticky routing, replay buffers and retention policy all went away, and what is left is the ordinary HTTP operations any team already runs, plus long-lived listen streams and the authorization story. Teams still serving 2025-11-25 clients pay both bills at once, which is the real cost of the dual-revision window. Teams writing new servers should mostly notice how much advice written about MCP operations in 2025 and early 2026 — the first version of this page included — is now advice about a transport that no longer exists.