MCP: Hosts, Clients, Servers

7 min read

P3
Deep Dive · Protocols & Interop

The Model Context Protocol: hosts, clients, servers, and three primitives.

MCP's participant model survived the 2026-07-28 revision intact; its lifecycle did not. There is no initialize handshake any more — every request restates its own protocol version and capabilities in params._meta, and an open connection is no longer a session. What is left is the part worth learning once: three roles (host, client, server), three intent-typed primitives (resources, tools, prompts), and a JSON-RPC layer where anything that has to persist between calls now needs a name you chose.

STEP 1

The participant model: host, client, server.

MCP — introduced by Anthropic in November 2024 and governed by an open specification — defines three roles. Getting them straight is the whole mental model:

  • Host. The application the user interacts with and that embeds the model — an IDE assistant, a desktop chat app, an agent runtime. The host manages the model, enforces user consent, and coordinates one or more clients.
  • Client. A connector inside the host, with a strict one-to-one relationship to a single server. If a host connects to three servers, it runs three clients. The client speaks the protocol and isolates that server's connection.
  • Server. A separate program that exposes capabilities — tools, resources, prompts — over the protocol. A server wraps one system (a filesystem, a database, a SaaS API) and is reusable by any MCP-capable host.
┌─ Host (the agent app) ─────────────────────┐
│  model + consent + orchestration            │
│   ├── Client A ───stdio───>  Server: files  │
│   ├── Client B ───HTTP────>  Server: github │
│   └── Client C ───HTTP────>  Server: db     │
└─────────────────────────────────────────────┘
   one client  ⇄  one server  (1:1, isolated)

The 1:1 client-to-server isolation is deliberate: it bounds trust per connection. A malicious or buggy server cannot see another server's traffic, and the host decides which servers a model context is exposed to. This is the protocol expression of the M+N argument from the interop-problem essay — wrap each system once as a server, teach the host the client once. One caveat that STEP 3 develops: a connection bounds trust, but since 2026-07-28 it holds no state, so "one client per server" is an isolation property and not a session.

STEP 2

Three server primitives: resources, tools, prompts.

An MCP server can offer three kinds of capability. The distinction is about who is in control:

Resources are application-controlled context: file-like data the host can read and place into the model's context — a file's contents, a database row, an API response. Each resource has a URI (for example file:///repo/README.md or a custom scheme). Resources are read-oriented and side-effect-free by intent; the host decides when to surface them.

Tools are model-controlled actions: functions the model may choose to invoke, each described with a JSON Schema input (the same substrate as the tool-calling-standards essay). Tools can have side effects — write a file, open a PR, run a query — so MCP expects host-mediated user consent before execution.

Prompts are user-controlled templates: reusable, parameterised prompt/workflow snippets a server publishes, which the host can surface as commands ("summarise this PR") and expand with arguments. They let a server ship expertise, not just raw capability.

Mnemonic for the control axis: resources = the application chooses what context to load; tools = the model chooses what action to take; prompts = the user chooses which workflow to run. Same protocol, three intents.

STEP 3

The wire: JSON-RPC 2.0, and the lifecycle that went away.

MCP messages are JSON-RPC 2.0: requests with an id expecting a response, results, errors, and one-way notifications. Through the 2025-11-25 revision every session began with an initialize handshake in which both sides exchanged a protocol version and a capabilities object. The 2026-07-28 revision deleted initialize and notifications/initialized and moved that negotiation onto every individual request: params._meta carries io.modelcontextprotocol/protocolVersion and io.modelcontextprotocol/clientCapabilities (both REQUIRED), io.modelcontextprotocol/clientInfo (SHOULD), and an optional io.modelcontextprotocol/logLevel. A request missing a required field gets -32602 and, over HTTP, status 400; a request that relies on a capability it did not declare gets MissingRequiredClientCapabilityError (-32021), again 400. The 2026-07-28 revision essay walks the full migration; what matters architecturally is the consequence — a request is now self-describing, and a connection carries no negotiated context.

The closest thing to a replacement is server/discover, and the asymmetry in its requirement level is the part to remember. Servers MUST implement it; clients are free never to call it, firing any RPC inline instead and handling UnsupportedProtocolVersionError (-32022) if the guess was wrong. When it is called, the response carries supportedVersions, capabilities, an optional instructions string, and ttlMs plus cacheScope so a client knows how long it may cache the answer and how widely. One detail has cost real deployments: serverInfo is not a top-level field of that result — it lives in _meta under io.modelcontextprotocol/serverInfo. The TypeScript SDK shipped the top-level shape and produced hard connect failures rather than graceful degradation, which is a durable reminder that "the SDK does it this way" is not the same claim as "the spec says so".

# 1. Optional for clients, mandatory for servers: what is this server?
{"jsonrpc":"2.0","id":1,"method":"server/discover",
 "params":{"_meta":{
   "io.modelcontextprotocol/protocolVersion":"2026-07-28",
   "io.modelcontextprotocol/clientCapabilities":{}}}}

# 2. Versions, capabilities, cache hint -- and serverInfo inside _meta.
{"jsonrpc":"2.0","id":1,"result":{
   "supportedVersions":["2026-07-28","2025-11-25"],
   "capabilities":{"tools":{"listChanged":true},
                    "resources":{},"prompts":{}},
   "instructions":"Search issues before opening a PR.",
   "resultType":"complete","ttlMs":3600000,"cacheScope":"public",
   "_meta":{"io.modelcontextprotocol/serverInfo":{"name":"github","version":"2.1"}}}}

# 3. Or skip it: every request restates the handshake for itself.
{"jsonrpc":"2.0","id":2,"method":"tools/list",
 "params":{"_meta":{
   "io.modelcontextprotocol/protocolVersion":"2026-07-28",
   "io.modelcontextprotocol/clientCapabilities":{},
   "io.modelcontextprotocol/clientInfo":{"name":"my-host","version":"1.0"}}}}

Two rules make that self-description survivable in a mixed fleet. Every result object now carries resultType — "complete", "input_required", or an extension value such as the Tasks extension's "task" — and a client MUST treat an absent resultType as "complete". That second rule is the entire backwards-compatibility story in one line: a server that has never heard of the field still returns results a current client reads correctly. And the state the handshake used to imply is gone with it, with no replacement. The spec requires that state spanning requests "MUST be referenced by an explicit identifier the client passes on each request", and forecloses the obvious workaround in one sentence: "an open connection, such as a STDIO process, is not a conversation or session." A stdio subprocess is a pipe; a held-open HTTP stream is a pipe. Anything that must persist between calls needs a handle you designed, returned in a result and threaded back as an argument.

The client discovers what a server offers with listing methods, then uses it with invocation methods:

server/discover   -> supportedVersions, capabilities, instructions, ttlMs
tools/list        -> [{name, description, inputSchema}, …]
tools/call        -> run a tool by name with arguments
resources/list    -> [{uri, name, mimeType}, …]
resources/read    -> fetch a resource's contents by uri
prompts/list      -> [{name, arguments}, …]
prompts/get       -> expand a prompt template with args
notifications/*   -> list_changed, progress, message, cancelled, …

Capability discovery is therefore runtime, not build-time: a host learns a server's tools by calling tools/list, or by calling server/discover once and caching the answer for ttlMs, and a list_changed notification tells it the set has moved. Those notifications are no longer pushed at a host unbidden — over HTTP a client asks for them by opening a subscriptions/listen stream and naming the ones it wants. Capability discovery has its own dedicated essay; the point here is that MCP puts it in the request path rather than in a lifecycle.

STEP 4

Transports, and where security enters.

MCP separates the message format (JSON-RPC) from the transport that carries it. Two transports are defined by the specification:

  • stdio. The host launches the server as a subprocess and exchanges JSON-RPC over its standard input/output. Ideal for local tools (a filesystem or Git server on your machine): no network surface, lifecycle tied to the process — though the open pipe is a transport and not a session, so it confers no permission to keep state on the side.
  • Streamable HTTP. The server is a remote HTTP endpoint; the client POSTs requests and the server answers with plain JSON or streams the response as server-sent events. Unsolicited server-to-client messages arrive on a subscriptions/listen stream the client opened deliberately — GET on the MCP endpoint, and the standalone stream it carried, were removed in 2026-07-28. This is the path for hosted, multi-client servers and is where authentication — the spec aligns remote auth with OAuth 2-style authorization — and transport security live.

One direction of travel was removed outright, and it is worth naming because older documentation is full of it. Through 2025-11-25 a server could call back into the host mid-execution: sampling/createMessage asked the host to run a model completion on the server's behalf, elicitation/create asked it to collect input from the user, roots/list asked which directories were in scope. 2026-07-28 removed all three as server-initiated requests. The capability survives as MRTR — Multi Round-Trip Requests — which keeps the feature and inverts who holds the call: instead of calling back, the server returns resultType: "input_required" with an inputRequests map and an opaque requestState, and stops. The client resolves those requests however it likes — asking the user, consulting a policy, filling a default — and re-sends the original method with inputResponses attached, under a different JSON-RPC id.

# The server needs a decision. It does not call back -- it returns.
{"jsonrpc":"2.0","id":7,"result":{
   "resultType":"input_required",
   "inputRequests":{
     "confirm":{"method":"elicitation/create",
                "params":{"message":"Merge PR 42 into main?",
                          "requestedSchema":{"type":"object",
                            "properties":{"ok":{"type":"boolean"}},
                            "required":["ok"]}}}},
   "requestState":"v1.opaque.9f2c81ae"}}

# The client answers, then re-sends the same call under a DIFFERENT id.
{"jsonrpc":"2.0","id":8,"method":"tools/call",
 "params":{"name":"merge_pr","arguments":{"number":42},
           "inputResponses":{"confirm":{"action":"accept",
                                        "content":{"ok":true}}},
           "requestState":"v1.opaque.9f2c81ae",
           "_meta":{"io.modelcontextprotocol/protocolVersion":"2026-07-28",
                     "io.modelcontextprotocol/clientCapabilities":{"elicitation":{}}}}}

The architectural payoff is that the control-flow inversion is gone. A server no longer needs a stream held open in order to reach the host mid-call, which is precisely why the session could be deleted at all — and the consent point moved somewhere better, because the host is now answering a question inside its own control flow rather than servicing an interrupt from a server it does not trust. Any documentation that describes a server "calling back into the host" is describing the superseded revision; the shape to look for now is a result that says it is not finished.

MCP standardises the channel, not trust. A connected server can return content that the model will read — a prompt-injection surface — and a tool call can have real side effects. The protocol's job is to make consent points explicit (the host gates tool execution and every input a server asks for mid-call); deciding what to grant is yours. Note too that clientInfo and serverInfo are self-reported and never verified by the protocol. The threat model, capability scoping, and provenance defenses are covered in the Safety & Agentic Security deep-dives. Treat "the server speaks MCP" as describing shape, never authorization.

The throughline: MCP is a participant model (host/client/server) plus three intent-typed primitives (resources/tools/prompts) plus a JSON-RPC layer whose negotiation rides on every request rather than on a lifecycle, over a pluggable transport (stdio or Streamable HTTP). That is the entire architecture; the rest of this track examines how its pieces — structured I/O, discovery, agent-to-agent extension — generalise.