Deleting the handshake deleted every fixture built on top of it — so the MCP tests worth writing now assert that each individual request describes itself, and the two in-process patterns everyone copied need different imports to still be true.
A suite written against the 2025-11-25 spec still passes, which is exactly the danger: it is green against a protocol your server no longer speaks. At 2026-07-28 there is no initialize, no Mcp-Session-Id and no SSE replay to resume, so session fixtures and resumability tests exercise nothing at all. The instincts survive intact — in-process beats subprocess, schema tests belong in a different file from behaviour tests, an agent loop is not a test — but the Python import moved, the TypeScript entry point changed outright, and the highest-value assertions available today (requestState integrity, the dual-era matrix, a foreign Origin) are ones almost nobody was writing a year ago.
The Python in-process test, and the two Clients that both work.
The biggest single lift in MCP testing is still the oldest advice on this page: do not spawn the server as a subprocess. Booting a Python interpreter and dialling a socket costs 200–500ms per test before anything useful happens; the in-process equivalent runs in single-digit milliseconds. That gap decides whether a fifty-case file gets run during development or only in CI, and a file that only runs in CI starts flaking — process boot racing a slow filesystem, a leftover port binding, a container with no spare core. Every second-order testing pathology (run only the changed test, trust only the ones that "usually pass", quarantine the rest) follows mechanically from a slow suite. The server-construction essay closes with a list of first-server mistakes; "wrote subprocess tests before trying the in-process client" belongs on it.
What changed is the door. There are now two correct Python Clients and they come from different projects. The official SDK, mcp v2.2.0, exports one at the top level — from mcp import Client — and it takes a server object directly. Two names in older tutorials are gone with it: mcp.shared.memory.create_connected_server_and_client_session has been removed, and mcp.server.fastmcp.FastMCP is renamed mcp.server.MCPServer. Separately, PrefectHQ's FastMCP 4.0.3 still ships fastmcp.client.Client unchanged, and it is still fine; note only that the repository moved from jlowin/fastmcp to PrefectHQ/fastmcp, so a bookmarked URL is not the maintained one. If a snippet you are copying does not say which package it imports from, it is ambiguous rather than wrong.
# tests/test_add.py — official SDK (mcp 2.2.0), in-process, no subprocess, no port import pytest from inline_snapshot import snapshot from mcp import Client from mcp.types import CallToolResult, TextContent from server import mcp @pytest.fixture def anyio_backend(): return "asyncio" @pytest.fixture async def client(): async with Client(mcp, raise_exceptions=True) as c: yield c @pytest.mark.anyio async def test_call_add_tool(client: Client): result = await client.call_tool("add", {"a": 1, "b": 2}) result.meta = None # drop the serverInfo stamp assert result == snapshot(CallToolResult( content=[TextContent(type="text", text="3")], structured_content={"result": 3}))
No mode= appears in that fixture, and it should not. Client(server) defaults to mode="auto", and because MCPServer answers server/discover on every transport including the in-process one, auto always lands on 2026-07-28 when it is talking to your own server. The three settings are "auto", "legacy", or a pinned version string; the pin is for the dual-era matrix in STEP 4, not for everyday use. Note also result.meta = None before the snapshot comparison: servers stamp io.modelcontextprotocol/serverInfo onto every result's _meta, and a snapshot that keeps it is a snapshot that breaks on every version bump.
raise_exceptions=True is worth understanding rather than cargo-culting, because it has a narrower meaning than its name suggests. It affects only failures outside a tool body — the ones production deliberately sanitises to "Internal server error" so that a stack trace never reaches a client. A test wants the real message, so leave it on. A failure inside a tool comes back as is_error=True either way, with or without the flag. And it has no meaning in production at all: it is a testing affordance, not a configuration knob you are choosing to leave unsafe.
TypeScript: handler.fetch, not createLinkedPair.
The TypeScript half of the old advice is the part that is now actively wrong. InMemoryTransport.createLinkedPair() has not been deleted — it is still exported, from @modelcontextprotocol/client — but the official documentation is unambiguous about its scope: "createLinkedPair connects 2025-era instances only; handler.fetch is the in-process entry for 2026-07-28 coverage." A linked pair therefore gives you a green suite that has never once exercised the revision you deploy. The packages to reach for are @modelcontextprotocol/server and @modelcontextprotocol/client at 2.0.0; v1 lives on as @modelcontextprotocol/sdk@1.30.0. The runner is vitest — if your MCP tests are still on jest, that is a second thing the 2.0 line assumes you have changed.
// tests/discount.test.ts — the transport never leaves the process
import { Client, StreamableHTTPClientTransport } from '@modelcontextprotocol/client';
import { createMcpHandler, McpServer } from '@modelcontextprotocol/server';
const handler = createMcpHandler(createServer); // createServer: () => McpServer
const transport = new StreamableHTTPClientTransport(new URL('http://test.local/mcp'), {
fetch: (url, init) => handler.fetch(new Request(url, init))
});
const client = new Client({ name: 'test-harness', version: '1.0.0' },
{ versionNegotiation: { mode: 'auto' } });
await client.connect(transport);
const result = await client.callTool({ name: 'apply-discount', arguments: { price: 80, percent: 25 } });
// teardown order matters
await client.close();
await handler.close();
The URL in that snippet is a label, not a destination: the transport never dials http://test.local/mcp — handler.fetch serves every request in-process, through the same createMcpHandler you deploy. That is the property the linked pair cannot give you, because the linked pair bypasses the handler entirely. Three follow-on details decide whether the harness stays honest. Teardown runs client-first, then handler, because handler.close() aborts any exchange still in flight, so a hung tool call cannot leak into the next test. A handler failure resolves as an ordinary result with isError: true, not a thrown error — there is nothing to catch, and a rejects.toThrow() assertion will simply never fire. And stdio has no in-process shortcut at all: if you ship a stdio server, that surface is tested over a real pipe or not at all.
NoBackChannelError is raised by your server, not by your test.
This is the most useful thing an in-process client will teach you, and it is easy to misread as a harness bug. NoBackChannelError is raised server-side, when server code reaches for a back-channel that a stateless connection does not have. Which side of the line your code falls on depends on how it asks:
- Resolver / dependency markers —
Elicit,Sample,ListRootsdeclared on a parameter. Works on a legacy connection (the SDK pushes the question); works on a 2026-07-28 connection too, where the SDK returns anInputRequiredResultinstead. - Imperative calls —
await ctx.elicit(...),ctx.session.create_message(),ctx.session.list_roots(). Works on a legacy connection; fails on 2026-07-28 — there is no back-channel to call.
The official wording is worth carrying verbatim, because it is also the migration argument: a resolver "works on every connection. For a client on a legacy connection the SDK sends it the question directly; on a 2026-07-28 connection the SDK returns the question from the call… Your resolver never knows the difference." So the guidance collapses to one line: migrate the server to resolvers and your test needs no mode= at all. Reach for mode="legacy" only when the server genuinely still calls ctx.elicit(), create_message() or list_roots(), or when you are testing a message_handler — and drop raise_exceptions=True in that fixture, because a legacy connection never sanitises in the first place and the flag buys you nothing.
Client-side the story is better than the era split suggests: one set of callbacks serves both eras. At 2026-07-28 the standalone server→client RPCs are gone, but the identical ElicitRequest, CreateMessageRequest and ListRootsRequest payloads ride inside input_requests and dispatch to the same callbacks you already wrote. Client retries with the answers and the echoed request_state and keeps going until a CallToolResult comes back; the intermediate rounds are invisible to the test body. Cap them with Client(..., input_required_max_rounds=10) so a server that keeps asking cannot hang the suite. Omit the callback entirely and call_tool raises MCPError("Elicitation not supported") — which is itself a perfectly good assertion. When you want to inspect the rounds rather than skip them, drop a level: client.session.call_tool(..., allow_input_required=True) and own the while isinstance(result, InputRequiredResult) loop yourself. The sampling and elicitation essay covers what those payloads mean; here they are just fixture inputs.
What a fixture sets up now: nothing — so assert the self-description instead.
There is no setup step left. In place of a negotiated session, every request must be self-describing, and that is precisely what a conformance-minded test now asserts. Required in params._meta: io.modelcontextprotocol/protocolVersion and io.modelcontextprotocol/clientCapabilities; io.modelcontextprotocol/clientInfo SHOULD be sent. Servers SHOULD stamp io.modelcontextprotocol/serverInfo onto every result's _meta — the field STEP 1 had to strip before snapshotting. The 2026-07-28 revision essay explains why the per-request restatement exists; what matters here is that it converts an invisible handshake into a list of checkable behaviours, each with its own code and HTTP status:
- Missing a required
_metafield →-32602, HTTP400. - Needs an undeclared capability →
-32021MissingRequiredClientCapabilityError, withdata.requiredCapabilitieslisting them, HTTP400. - Header/body version disagreement, or a missing standard header →
-32020HeaderMismatch, HTTP400. Required on every POST:MCP-Protocol-Version,Mcp-Method, andMcp-Namefortools/call,resources/readandprompts/get. Values may arrive Base64-sentinel-encoded as=?base64?…?=, and servers MUST decode before comparing — send one that way and see whether yours does. - Unsupported version →
-32022, withdata.supportedlisting the versions. - Unknown method over HTTP →
404plus-32601. - A notification POST →
202 Accepted, no body. - Legacy traffic →
GETandDELETEanswer405; anMcp-Session-Idheader is ignored, neither minted nor echoed;Last-Event-IDis ignored. - Bad
Origin→403 Forbidden.
The other half of the update is deletion, and it is larger than the addition. Session fixtures, Mcp-Session-Id plumbing, initialize/initialized ordering assertions, DELETE-to-terminate teardown, and every SSE resumability test — Last-Event-ID handling, event-id monotonicity, replay-on-reconnect — should come out of the suite entirely. They are not merely redundant; they assert behaviour a conformant server MUST NOT exhibit. A broken stream at 2026-07-28 has one correct handling: the client re-issues the work as a new request with a new id.
A stateless modern handler is constructed per request. A fixture that mutates server state between two calls — seeding a cache, flipping a feature flag on the server object, stashing a counter — silently does nothing, because the second call gets a fresh instance. The test still passes, for the wrong reason, and it will keep passing after the behaviour it was meant to protect is gone.
What is genuinely new is a set of assertions almost nobody wrote before, because before the revision there was nothing to assert. These are cheap, mechanical, and they belong in the schema-contract file rather than the behaviour file — the same split the schemas and contracts deep-dive argues for in general, and the reason to keep them separate is unchanged: a contract break hits every client at once, including ones you do not own.
server/discoveris implemented. Servers MUST implement it; a surprising number of ported servers do not.tools/listdoes not vary per connection (it MAY vary by the authorization presented). Test it directly: list, call some tools, list again, assert the two manifests are identical.ttlMsandcacheScopeare present ontools/list,prompts/list,resources/list,resources/readandresources/templates/list.- Tools come back in a deterministic order (SHOULD) — a dict-iteration-order manifest is a snapshot test that flakes for no reason.
- A state handle is not authentication. Servers MUST NOT treat possession of a handle as authentication and SHOULD key state as
<user_id>:<handle>. Mint one as principal A, present it as principal B, assert refusal. - The dual-era matrix. Modern↔Modern works; Modern client → Legacy server fails; Legacy client → Modern server fails; a dual-era server works both ways. Test only the cells you actually claim to support.
resultType, requestState integrity, and the auth seams.
Every result MUST carry resultType, either "complete" or "input_required". An unrecognised value MUST be treated as invalid; an absent one MUST be treated as "complete", which is the rule keeping older servers working. Assert it at the wire level, not through a typed client — the typed client has already normalised the field, so a server that omits it looks identical to one that sets it correctly, right up until a stricter client meets it.
The retry rule is the one implementations break, and the spec states it plainly: "Note that the JSON-RPC id MUST be different between the initial request and the retry." The shape below is worth pinning in a fixture because the nesting is easy to get wrong — inputResponses and requestState sit directly inside params, alongside arguments, not inside it.
// round 1 response
{"jsonrpc":"2.0","id":2,"result":{
"resultType":"input_required",
"inputRequests":{"github_login":{"method":"elicitation/create","params":{...}}},
"requestState":"eyJsb2NhdGlvbiI6Ik5ldyBZb3JrIn0..."}}
// round 2 request — NEW id
{"jsonrpc":"2.0","id":3,"method":"tools/call","params":{
"name":"get_weather","arguments":{"location":"New York"},
"inputResponses":{"github_login":{"action":"accept","content":{"name":"octocat"}}},
"requestState":"eyJsb2NhdGlvbiI6Ik5ldyBZb3JrIn0..."}}
// tamper: flip one byte of requestState, retry, assert the frozen error
{"code": -32602, "message": "Invalid or expired requestState"}
requestState integrity is a MUST, and it is the assertion most missing from MCP suites. The client holds that state between legs, which means what comes back is client-supplied input: modifiable, expirable, or lifted wholesale from a different call. The test is three lines. Run one round, flip a byte in requestState, retry, and assert the frozen error {"code": -32602, "message": "Invalid or expired requestState"} — one message for every cause, so the wire never reveals which check failed. Then repeat it twice more: once replaying a valid state captured from a different call, once presenting a valid state minted by a different principal. All three must give the same answer. A server that distinguishes them in its error text has built an oracle.
There is an asymmetry here worth stating plainly, because it silently decides whether that test passes for free or fails on day one. Python's MCPServer seals requestState by default, under a key generated at process start; configure RequestStateSecurity(keys=[...]) with keys of at least 32 bytes for any deployment that spans instances or must survive a restart, and note the TTL defaults to 600s and binds the authenticated principal. TypeScript does not seal by default. You opt in with createRequestStateCodec from @modelcontextprotocol/server, whose key MUST be at least 32 bytes or construction throws RangeError. Python's low-level Server is likewise unsealed. And no SDK can know which of your questions a given answer belongs to — include your own question identifier in the state and check it on retry. One more era guard: returning an InputRequiredResult on a legacy connection yields -32603 "Handler returned an invalid result", so a dual-era server must branch on the protocol version before it decides how to ask.
Authorization is testable without standing up an identity provider, which is the excuse most often given for not testing it. TokenVerifier is a protocol with exactly one async method — verify_token takes the raw token and returns an AccessToken or None, and there is nothing else to implement — so stub it and drive every branch from the test. TypeScript exposes the same seams: verifyBearerToken, requireBearerAuth, buildOAuthProtectedResourceMetadata. The trap is that get_access_token() returns None in-memory and over stdio, so any per-scope authorization logic has to be tested over real HTTP — the one place the in-process harness genuinely cannot reach. The OAuth 2.1 profile essay covers what those tokens must contain.
Conformance, the Inspector, and what to stop recommending.
There is now an official conformance suite, @modelcontextprotocol/conformance, and a version trap on the way in: npm latest is 0.1.16, but the SDKs themselves run 0.2.0-alpha.11. Prefer the pinned npx form over the official composite Action, whose tags stop at v0.1.16.
$ npx --yes @modelcontextprotocol/conformance@0.2.0-alpha.11 server \
--url http://localhost:3001/mcp --requirements 2026-07-28 \
--expected-failures ./conformance/baseline.yml
$ npx @modelcontextprotocol/inspector --cli node build/index.js \
--method tools/list --strict --format json \
| jq -e '[.schemaFindings[]?.findings[]? | select(.severity=="error")] | length == 0'
The suite runs per-scenario assertions and validates every JSON-RPC message against the spec JSON Schema, which is more than a hand-written harness will do. The frozen 2026-07-28 set is 37 required server scenarios, 32 client, 20 unscored — including 11 MRTR scenarios plus server-stateless, dns-rebinding-protection and caching. Two operational details decide whether it tells you the truth. Scenarios run at their own revision's wire version, so a scenario that applies to both revisions must be run once under each; one run does not cover the other. And the expected-failures baseline ratchets both ways: a failure listed in the baseline exits 0, a new failure exits 1, and a scenario that starts passing while still listed also exits 1, which is what stops the baseline rotting into a permanent excuse. One gap worth flagging before you rely on it: the auth/ scenarios exist only for clients. If you run a protected server, conformance does not test your token validation at all.
The claim that the Inspector is a debugger and not a test has been obsolete since v2. Inspector v2.6.0 ships a --cli client built for CI, with nine stable exit codes: 0 success, 1 usage or unexpected, 2 no MCP App found, 3 server requires authentication, 4 unreachable, 5 tool error, 6 --strict schema-portability error, 7 --verify skills violation, 8 --verify incomplete. On any non-zero it writes a one-line JSON error envelope to stderr — take the last line. --strict is the part that is real contract testing, and its rationale is the quotable bit: a census of 617 public servers found zero that fail the SDK's own parser. Plain JSON Schema validation finds nothing; what actually breaks clients is the narrower subset each consumer accepts. Gotcha: the Inspector's protocol era defaults to legacy and is a config-file field only — protocolEra, with no CLI flag — so a --cli run exercises the legacy path unless you pass a config setting "protocolEra": "modern".
Against that, the tool this page used to recommend is effectively dead. Microsoft's mcp-interviewer last released v0.0.12 on 2025-10-11, its commit log after that is Dependabot-only and stops on 2025-12-01, and Microsoft's own README describes it as research/experimental. Do not build a CI gate on it. The same verdict applies across most of the third-party testing and scanning layer: steviec/mcp-server-tester, mclenhard/mcp-evals, f/mcptools, PyPI mcp-testing-framework, and the scanner cluster of mcp-shield, mcpSafetyScanner, mcp-guardian and mcp-context-protector are all unmaintained; Janix-ai/mcp-validator, which circulated widely in 2025 link lists, is now a 404. The maintained surface is the conformance suite, the Inspector, and the tests you write yourself.
Wiring conformance into CI has a shape both official SDKs converged on, and every line of it is a scar. Precheck the port and refuse to start if something is already listening, because a readiness check cannot tell your server from a stale one. Use kill -0 liveness inside the wait loop so a crashed server fails in a second instead of burning the full timeout. Give curl a --max-time so a black-holed listener cannot wedge the loop forever. And trap the cleanup so a failed assertion does not leak a process into the next job.
# ci/conformance.sh — precheck, liveness, bounded curl, trapped cleanup
set -euo pipefail
lsof -i :3001 -sTCP:LISTEN -t >/dev/null && { echo "port 3001 already in use"; exit 1; }
node build/index.js --port 3001 --stateless &
SERVER_PID=$!
trap 'kill "$SERVER_PID" 2>/dev/null || true' EXIT
for _ in $(seq 1 50); do
kill -0 "$SERVER_PID" 2>/dev/null || { echo "server exited during startup"; exit 1; }
curl -fsS --max-time 2 http://localhost:3001/health >/dev/null && break
sleep 0.2
done
npx --yes @modelcontextprotocol/conformance@0.2.0-alpha.11 server \
--url http://localhost:3001/mcp --requirements 2026-07-28
Run both eras in one job by starting the server twice — once stateful against --requirements 2025-11-25, once stateless against --requirements 2026-07-28 — rather than trusting one run to imply the other. A reality check is fair here: none of the flagship MCP servers currently run the official conformance suite. Adoption lives in the SDKs and the gateways, which means running it puts you ahead of the servers you are interoperating with, not merely level with them.
If you adopt nothing else on this page, adopt three tests whose coverage is wildly out of proportion to their cost. First, send a foreign Origin and assert 403. Four lines, and it covers the single most repeated MCP vulnerability: DNS-rebinding and missing origin validation shipped as seven CVEs across six SDKs plus the Inspector itself, including CVE-2025-49596 at CVSS 9.4 — the pattern catalogued in the security anti-patterns essay. Second, fire N concurrent tools/calls, once with unique JSON-RPC ids and once with duplicates. That catches CVE-2026-25536, a shared transport leaking one client's response to another and "most common in stateless deployments", and an open python-sdk issue where two concurrent requests sharing an id cross-wire — whose reporter observed one user's tool response delivered to a different conversation's request in production. Third, snapshot tools/list with a byte ceiling, then re-check it after several calls. The re-check is the whole point: the Deadbugz campaign mutates tool descriptions only after the third call, so a single-shot assertion passes cleanly — see tool poisoning for the mechanism. The same test catches schema-dialect breaks and context-blowing catalogues on the way past.
Two negatives are worth stating rather than hiding, because pretending otherwise is how a suite acquires theatre. There is no healthy load-testing option for MCP today; if you need throughput numbers you will be building the harness yourself, and you should budget for that rather than shop for it. And prompt-injection incidents are not server-test-catchable: no server-side assertion fails, because the server did exactly what it was asked. You test the architectural control — the allowlist, the confirmation gate, the scoped token — not the injection. That is also the honest version of the old advice about agent loops. Exploratory agent runs are how you find out what to test, and they remain worth doing; they are not the thing that tells you a contract held, because a competent agent reads the error, retries with the corrected argument name, and produces an output that looks fine over a schema regression that will break every cached client you do not own.