Subagents

A31
Concepts · Agentic AI Explained

Subagents.

A subagent buys you exactly one thing — a second context window that fills up and then vanishes — and most teams pay for it expecting a colleague. The persona is free and does nothing; the isolation is the product. Everything the child read is gone the moment it returns, so what you actually ship is the paragraph it wrote on the way out. The return channel, not the role, is the design.

STEP 1

A subagent is a context boundary with a task string on one side.

Strip the vocabulary away and the mechanism is small. The parent sends a task description; a fresh model instance starts with an empty (or near-empty) context, a tool set the parent chose, and a step budget; it works; it returns text. The parent's context gains that text and nothing else. No shared memory, no back-and-forth, no state that survives.

That shape has three properties worth paying for, and they are the only three:

  • Context isolation. The child's forty tool results, its dead ends, and its 90,000 tokens of grep output land in its window, not yours. This is the property people actually want, and it is why subagents feel like they make long tasks tractable — see context budgeting.
  • Tool scoping. A child that was handed read-only tools cannot write, whatever it concludes it should do. That is a real containment boundary and it composes with blast radius thinking.
  • Parallelism. Independent children run at once, which converts a serial chain of context-consuming work into wall-clock you can afford.

Two neighbours to keep separate. A multi-agent system has peers that exchange messages and hold state across turns; a subagent is a one-shot call that happens to be answered by a model. And a tool call is deterministic — you know what comes back. A subagent is a tool call whose return value was written by a language model, which is a different thing to depend on.

STEP 2

The persona is not the mechanism, and naming one costs you nothing but buys you nothing.

The common design is a cast: a "researcher", a "critic", a "planner", a "security reviewer". It reads like organisational design, which is why it is persuasive. But a critic subagent and a critic paragraph in the same prompt differ in exactly one respect — the subagent does not see the conversation that produced the thing it is criticising. Sometimes that independence is the point, and you should say so. Usually it is an accident, and the critic reviews a summary of the work rather than the work.

The test that actually predicts whether a subagent pays is a ratio: how much context does it consume, and how much does it return?

  • High ratio — delegate. "Find every call site of this function and tell me which ones pass a nullable argument" reads thousands of lines and returns a list of six. Searching, log triage, scanning a directory, reading a long document to answer one question: all large-in, small-out.
  • Low ratio — inline it. If the child returns nearly everything it saw, you paid a full prompt, a round trip, and a summarisation step to move tokens from one window into another. Rewriting a paragraph, making a judgement call on material the parent already holds, formatting output — none of these are subagent work.
  • Ratio of one — you built a router. A subagent that takes a question and returns an answer with no intermediate work is a model call with extra latency. That may be fine as a routing step; it is not delegation.

Apply the ratio before the org chart. Most cast lists survive it in two roles and lose the rest.

STEP 3

The summary is the entire deliverable, and the parent cannot check it.

This is the part that gets designed last and causes the most trouble. When the child returns, its evidence is unrecoverable: the tool results are not in the parent's context, the reasoning that weighed them is gone, and the parent has no cheap way to look again. What arrives is prose, asserted with whatever confidence the child happened to use, and the parent will treat it as fact — because in its context it is a fact, indistinguishable from a tool result.

Two failures follow directly, and both are quiet:

  • "I found nothing" and "I failed to look" produce the same sentence. A child whose search tool errored, whose budget ran out, or which searched the wrong directory reports an empty result in the same words as a child that searched correctly. The parent proceeds on a negative it has no reason to doubt.
  • Detail gets invented in the compression. The child summarises twelve files into four sentences, and the fourth sentence contains a line number that is approximately right. This is ordinary grounding failure, but it is worse here, because the evidence that would contradict it has already been discarded.

The fix is to stop treating the return value as a report and start treating it as a record:

  • Return handles, not retellings. File paths, line ranges, record IDs, URLs — things the parent can re-open for a few hundred tokens if it needs to. A summary that cites its sources by handle converts an unverifiable claim into a cheap lookup. This is the same discipline as shaping tool results.
  • Make the shape structured and make absence explicit. A schema with searched, found, and incomplete_because fields forces the child to distinguish the three outcomes that prose collapses into one. A child that hit its step limit must be unable to report a clean negative.
  • Carry provenance into the parent. The parent's context should record that a claim came from a subagent, not from its own observation, so that a later contradiction is resolvable rather than confusing — the argument instruction hierarchy makes about trust levels applies to your own children too.
STEP 4

Delegation is not free, and the costs land somewhere you are not looking.

The economics are worth stating plainly, because the intuition that a subagent "saves context" quietly implies that it saves money, and it does not.

  • Every child is a full prompt. System prompt, tool schemas, task description — paid again, per child. A fan-out of five is five of those, and the parent then pays to read five summaries. You are trading total tokens for parent-context tokens, and that is usually the right trade; it is still a trade, and it belongs in your cost model explicitly.
  • Parallel only helps if the work is independent. Children that must see each other's findings are not parallel; they are a sequence you have made harder to debug. If the second child's task depends on the first child's answer, the wall-clock win is zero and the context cost doubled — the same constraint as parallel tool calls.
  • The goal reaches the child only through the task string. Whatever the parent understood and did not write down is not transmitted. Each delegation level is another lossy restatement of the objective, which makes goal drift compound with depth rather than stay constant.
  • Nesting makes the trace a tree nobody reads. Children that spawn children turn a linear transcript into something your observability stack probably renders as unrelated runs. When the result is wrong, you now have a credit assignment problem instead of a transcript to read. Cap the depth at one until you can see two.

Before adding a subagent, answer two questions in writing: what large thing does it read that the parent must not, and what small structured thing does it return. If you cannot answer the first, inline the prompt. If you cannot answer the second, you have not finished designing it — and the default free-text return is the one that will fail silently in production. Then go read the last twenty summaries your subagents produced and check how many claims in them you could verify without re-running the child; that fraction is the real quality of your delegation layer.

Related: sub-agent patterns compared for how the major frameworks implement this, single vs multi-agent for when to stop at one, and context compaction for the other way to spend a window you have run out of.