Sampling & Elicitation via MRTR

9 min read

C7
Deep Dive · MCP

A server still has to ask the host for a model completion, a user's answer, or a filesystem root — but since the 2026-07-28 revision it can no longer call back to get them: it returns the questions as a result, and the client retries the original call carrying the answers.

Every sampling and elicitation tutorial written before mid-2026 describes a mechanism the protocol no longer has. The 2026-07-28 revision deleted server-initiated requests outright — the spec's words are "This is a breaking change" — and replaced them with MRTR, Multi Round-Trip Requests: the server answers with resultType: "input_required" and its questions attached, then the client re-sends the same method with the answers. The reasons a server needs to ask the host things survive untouched. The mechanism is inverted.

STEP 1

What a server still needs from the host — and what stopped working.

Three things a server cannot produce on its own still have to come from the other end of the connection, and none of the reasons changed. A server that wants an LLM but does not want to become an LLM operator needs a completion run on the host's account: summarize a resource before returning it, classify a user-supplied blob against a taxonomy the tool understands, rewrite a diff into a commit message the calling agent will then decide whether to apply. Solve any of those server-side and the author buys a model provider account, an API key, a bill, and a compliance conversation. A server that is missing one structured fact — which repository, which of three ambiguous matches, confirm-before-destroying — needs to reach the user without routing the question through the model and losing the answer in a game of telephone. And a server that operates on files needs to know which directories the user has actually put on the table. The MCP participant model explains why all three point the same direction: the API key, the user's attention, and the consented filesystem are the host's assets, not the server's.

What stopped working is the delivery. Until 2026-07-28 the server issued sampling/createMessage, elicitation/create or roots/list as a request of its own, back down the session it was already serving. That required a bidirectional stream held open across the whole operation, which in turn forced sticky routing and shared state: two round-trips of one logical call had to land on the same server instance, with that instance's memory of the call still warm. The spec now says the opposite. "Servers MUST send server-to-client requests (such as roots/list, sampling/createMessage, or elicitation/create) using the MRTR pattern. The previous pattern of server-initiated requests is no longer supported. This is a breaking change." MRTR exists because the sticky-instance requirement had to go; it shipped alongside statelessness for exactly that reason. The 2026-07-28 revision essay is the fuller account of what else moved at the same time, and Streamable HTTP covers the transport-level bill this was paying down.

STEP 2

MRTR: the server returns its questions, the client retries the call.

MRTR is permitted on exactly three client requests — prompts/get, resources/read, and tools/call — and servers MUST NOT send an InputRequiredResult in response to any other client request. In place of the normal result, the server returns resultType: "input_required", an inputRequests map whose keys it assigns itself, and an opaque requestState string. Values in inputRequests MUST be one of ElicitRequest, CreateMessageRequest, or ListRootsRequest — the same three payloads the old callbacks carried, now travelling as cargo inside a response instead of as requests in their own right. A server MUST include at least one of inputRequests or requestState; state with no questions is the load-shedding case, the server saying "come back with this and I will carry on" without needing anything from the user at all. And a server MUST NOT put an entry in inputRequests for a capability the client never declared, which makes the client's capability declaration a hard gate rather than a hint.

// client -> server : an ordinary call
{ "jsonrpc": "2.0", "id": 1, "method": "tools/call",
  "params": { "name": "open_pr", "arguments": { "title": "Fix retry backoff" } } }

// server -> client : not an answer — a set of questions, plus its own continuation
{
  "jsonrpc": "2.0",
  "id": 1,
  "result": {
    "resultType": "input_required",
    "inputRequests": {
      "github_login": {
        "method": "elicitation/create",
        "params": {
          "mode": "form",
          "message": "Please provide your GitHub username",
          "requestedSchema": {
            "type": "object",
            "properties": { "name": { "type": "string" } },
            "required": ["name"]
          }
        }
      },
      "capital_of_france": {
        "method": "sampling/createMessage",
        "params": {
          "messages": [
            { "role": "user", "content": { "type": "text", "text": "What is the capital of France?" } }
          ],
          "maxTokens": 100
        }
      }
    },
    "requestState": "eyJsb2NhdGlvbiI6Ik5ldyBZb3JrIn0"
  }
}

The client fulfils each entry locally — renders the form, runs the completion, enumerates its roots — and then re-sends the original method with an inputResponses map keyed identically to inputRequests, plus the state it was given. Two rules on that retry are where implementations break. First, the JSON-RPC id MUST differ between the initial request and the retry: these are two independent requests that happen to be about the same operation, not one request resumed, and the instinct to reuse id: 1 because it "feels like a continuation" is the single most common mistake in MRTR client code. Second, the client MUST echo requestState byte-exact and MUST NOT inspect, parse or modify it — and if the server did not send one, the client MUST NOT invent one. The correlation that used to live in a held-open connection now lives entirely in those two maps and that one opaque string.

// client -> server : same method, NEW id, answers keyed to the server's own keys
{
  "jsonrpc": "2.0",
  "id": 2,
  "method": "tools/call",
  "params": {
    "name": "open_pr",
    "arguments": { "title": "Fix retry backoff" },
    "inputResponses": {
      "github_login": { "action": "accept", "content": { "name": "octocat" } },
      "capital_of_france": {
        "role": "assistant",
        "content": { "type": "text", "text": "The capital of France is Paris." },
        "model": "claude-3-sonnet-20240307",
        "stopReason": "endTurn"
      }
    },
    "requestState": "eyJsb2NhdGlvbiI6Ik5ldyBZb3JrIn0"
  }
}
STEP 3

Elicitation, form and URL, with the callback removed.

Elicitation is the one of the three that is not deprecated, so it is the one worth building on. The form-mode payload itself is unchanged: a JSON Schema describing the fields the server wants, a human-readable message explaining why, and a host free to render checkboxes for booleans, dropdowns for enums, and text inputs for strings. The two design rules that survived contact with practice survive this revision too. Keep the schema flat and small — no nested objects, three or four fields at most — because a dialog with fifteen inputs stops a workflow harder than an extra round-trip would. And give every optional field a default, so the user can accept and move on without typing. What changed is that the form arrives as a value inside inputRequests rather than as a request the server sent, so the server is not sitting on a half-finished call while the user decides; it has already returned, and it will hear back only if the client retries.

Two pieces of 2025-11-25 machinery went away with the callback. notifications/elicitation/complete and the elicitationId that went with it are gone, because there is no longer an out-of-band completion signal to correlate — the retry is the correlation. notifications/roots/list_changed is gone as well. If your server or client still sends either, it is speaking a revision that no longer exists.

URL-mode elicitation is the variant the OAuth 2.1 profile essay points at every time it explains what to do instead of token passthrough, and MRTR makes its shape more honest rather than less. The tool call needs to act on a third-party service the server holds no credentials for. The server returns an ElicitRequest in url mode alongside its requestState and then stops — no open connection, no parked thread, nothing to keep warm. The client's host opens the URL in the user's browser, the user completes the flow, and the third-party token lands on a callback the client controls. Only when the client retries does the server learn anything, and what it learns is an accept action and whatever handle it needs to identify the linked account — never the third-party token. Token passthrough is forbidden because the naive version, where the server holds the caller's token and forwards it downstream, collapses the RFC 8707 audience guarantee and forces the downstream service to trust the wrong principal. The non-negotiable is unchanged: the third-party token never touches the server.

t=0.00  client  | tools/call id=1                 | name=connect_github
t=0.01  server  | result resultType=input_required | inputRequests.link = elicitation/create mode=url
t=0.01  server  | requestState=AEAD(...)           | request COMPLETE — no connection held, no state parked
t=0.02  client  | opens url in user's browser      | user completes OAuth
t=8.31  client  | callback lands on client         | code -> exchange -> access_token STORED IN CLIENT
t=8.33  client  | tools/call id=2                  | inputResponses.link={action:accept}  requestState echoed byte-exact
t=8.34  server  | verify requestState, then parse  | resumes via linked account | server never sees gh access_token
STEP 4

requestState is your continuation, handed to an untrusted party.

This is the part of MRTR that has no analogue in the old design, and it is the part to get right. When the connection held the state, the state was in the server's own memory. Now the server serializes its continuation, hands it to the client, and asks for it back — and the spec is blunt about what that means: servers MUST treat requestState as attacker-controlled. If it influences authorization or business logic in any way, the server MUST integrity-protect it with an HMAC or an AEAD, and MUST reject state that fails verification. It SHOULD embed the principal, a TTL, and an identifier for the originating request. A server that stuffs {"tenant": "acme", "approved": true} into requestState unsigned has not built a resumption token; it has built an authorization decision that the caller can edit, and the caller is exactly the party the decision was meant to constrain. The MCP security anti-patterns essay catalogues what happens when a server trusts the wrong end of a message; this is the newest way to arrive there.

// what belongs inside requestState, before you seal it
{
  "sub":  "user_9x",                     // principal the ORIGINAL request ran as
  "orig": "tools/call:open_pr:01JB7Q",   // originating-request identifier
  "exp":  1793664000,                    // TTL — minutes, not days
  "step": "awaiting_github_login"        // the server's own continuation
}

// on the wire : AEAD(key, canonical_json(state))  ->  opaque base64
// on return   : VERIFY first, then parse. Verification failure is a hard reject,
//               not a warning. Then check exp, then check that sub still matches
//               the principal on the retry — a valid blob from another session
//               is still the wrong blob.

The operational tells are short. Bind the state to the principal on the retry, not just to the principal that created it, or you have built a replay primitive that works across sessions. Keep the TTL measured in minutes, because the only legitimate gap between the two round-trips is however long a human takes to fill in a form or finish an OAuth flow. And cap retries per logical operation on the client side: a server that answers every retry with another input_required is either malfunctioning or walking the client around a loop, and the client is the only participant positioned to stop it.

STEP 5

The mechanism is mandatory; two of the three payloads are on a clock.

The most useful sentence on this page for anyone choosing what to build on: MRTR is how you must carry these requests today, and two of the three requests it can carry are themselves deprecated. Sampling, Roots and Logging are all marked Deprecated in 2026-07-28. Elicitation is not. Sampling-with-tools — the pattern where the server declares its own tools inside the nested completion and borrows the host's whole agent loop rather than just a model call — is part of Sampling and carries the same marker; so does includeContext: "thisServer" and "allServers", deprecated back in 2025-11-25 and now following the feature that hosts it. Read that as a planning constraint, not a countdown: put the user-input path on elicitation, and treat sampling as a convenience your server should degrade gracefully without.

Be precise about what the deprecation policy actually promises, because the internet is already getting it wrong. The earliest a deprecated feature may be removed is "the first revision released on or after 2027-07-28". That date marks when a feature becomes eligible for removal, not when it is removed — the actual removal is a Core Maintainer decision taken at release preparation, and it may happen later or not in that cycle at all. Nothing has been removed under this policy yet. "Sampling is removed in July 2027" is a claim nobody can currently make, and a roadmap built on it will be wrong in whichever direction the maintainers go.

The consent story survives the inversion and gets slightly better. Sampling and elicitation still let a server drive an experience the user did not ask for at that moment, which still makes consent a per-interaction question rather than a per-connection one — the discipline the human-in-the-loop essay treats as an ops problem. What MRTR adds is batching: because the server returns all its questions in one inputRequests map, the client can render one consent surface for the whole set instead of interrupting the user once per callback. That surface still has to show which server is asking, why, what the model would be shown for a sampling entry, and — for a URL-mode elicitation — the origin of the URL about to open in the user's browser. Keep the three grants distinct: once for this call, for this session, forever for this server. A user who approved one summarization has not approved the next hundred. Both primitives work exactly as well as the client that renders them, and both fall over as hard as any web permission dialog when the client renders them lazily.