Parallel Tool Calls

B27
Concepts · Core Building Blocks

Parallel Tool Calls.

The batch your model just emitted was decided before any of it ran — four calls chosen in one breath, with no result from the first available to shape the second. That single property turns parallel tool calling from a latency optimisation into a correctness decision: the only calls safe to issue together are the ones that commute and can be retried independently, and the model is not the component that knows which of yours those are. You are, and the place to say so is the executor, not the prompt.

STEP 1

One turn, several calls, no lookahead.

In an ordinary tool-calling turn the model emits one call, your harness runs it, and the result goes back as the next message. A parallel batch is the same protocol with the count raised: a single assistant turn carries several tool-use blocks, the harness executes them — usually concurrently — and returns every result together before the model speaks again.

The interesting part is not the concurrency. It is that all N calls were selected from the same information. In a sequential loop, call two is written after result one exists and can be narrowed by it. In a batch, call two is a guess made in parallel with call one, and the model does not learn it guessed until the whole batch comes back.

  • It is on by default. Both major APIs emit multiple calls per turn unless told otherwise: OpenAI takes a top-level parallel_tool_calls: false, Anthropic takes tool_choice: {"type": "auto", "disable_parallel_tool_use": true}, which caps the turn at one tool. If you have never set either, you already ship batches.
  • The flag is per request, not per deployment. That is the useful shape: a research phase can fan out while the phase that writes anything runs one call at a time, inside the same run.
  • Concurrency is your harness's choice. The API hands you a list. Nothing obliges you to execute it concurrently, and for a mixed batch you should not.

Say the batch out loud and the risk becomes obvious: "run these four things in an order I have not specified, and tell me what happened afterwards." No transaction, no ordering guarantee, no chance to abort after the first result. That is a fine instruction for four lookups and a bad one for anything that writes.

STEP 2

Two properties decide whether a batch is safe.

A call belongs in a batch when it satisfies both of these against every other call in the same batch. One is not enough.

  • Commutativity. Running A then B leaves the world in the same state as B then A. Four reads against a store nobody is writing to commute. apply_discount and charge_card against the same order do not, and neither does a read of a record paired with a write to it — the read may or may not see the write, and which it sees is a scheduling accident.
  • Independent retryability. If one call fails, re-running only that one is correct. This is where batches usually break, because partial failure is the ordinary outcome: three succeed, one returns a 503, and the model — which sees a mixed result set and no repair protocol — commonly re-emits something close to the original batch. Any non-idempotent call in it now runs twice.
# A batch the model will happily emit, and should not

assistant turn:
  create_ticket(title="refund request")      → ok, id=T-4192
  post_comment(ticket="?", body=…)           → error: needs T-4192
  notify_customer(ticket="?")                → error: needs T-4192

# The model retries the batch. create_ticket is not idempotent.
# You now have T-4192 and T-4193, and a customer notified about one of them.

The failure is not that the model was careless. It is that the batch shape gave it no way to be careful: the dependency it needed to respect was only expressible by waiting, and waiting is the thing a batch does not do.

STEP 3

What you actually buy, and what it quietly costs.

The speedup is real. It is also narrower, and more expensive elsewhere, than the pitch suggests.

  • Latency becomes max, not sum — which pays when the calls are slow and roughly equal, and pays almost nothing when one call dominates. Three 80 ms lookups behind a 4 s search save you 240 ms on a turn that still takes four seconds.
  • Tokens go the other way. Every result lands in context at once, in full. A sequential loop can stop after result one made results two through five unnecessary; a batch has already paid for all five, and paid again on every subsequent turn that carries them. Fan-out is the fastest way to fill a context window with material the run never needed, which is why it shows up as a cost problem before it shows up as a latency win.
  • It does more work, not the same work faster. A model that can see result one often issues a narrower second call, or none. Batching removes that option by construction, so the comparison is not "same four calls, less wall clock" — it is "four calls instead of the two a sequential loop would have made".
  • Spend gets lumpy. Concurrency multiplies against rate limits and per-call cost in the same instant, so a run that was comfortably inside a quota sequentially can trip it in a batch, and the retry storm that follows is charged to you.

The honest rule of thumb: batch when the calls are independent reads that are individually slow — search, retrieval, a fan-out over several sources — and serialise everything else. That is a narrower licence than the default setting gives you.

STEP 4

Put the constraint in the executor, not the prompt.

"Do not call these two tools together" is an instruction with a failure rate. The same rule in the code that dispatches the batch has a return value. Four changes, all in your harness, none requiring the model's cooperation:

  • Classify every tool as reading or mutating in the registry, next to its schema. This is one boolean per tool definition and it is the field the rest of this list depends on.
  • Refuse the unsafe shape at dispatch. A batch containing more than one mutating call gets serialised — or rejected with a message telling the model to issue them one at a time, which is a repair it handles well. A batch mixing a write with a read of the same resource gets the same treatment.
  • Derive idempotency keys from arguments, not from a fresh UUID. A key generated per attempt is not an idempotency key. Hash the call's semantic arguments plus a run identifier so a re-emitted create_ticket resolves to the ticket that already exists. The operational depth is in idempotency and retries.
  • Return per-call results with stable identifiers and per-call errors — never one merged blob, and never a single failure that discards three successes. The model can only retry precisely if the transcript says precisely what happened; see shaping tool results.
  • Cap the width. An upper bound on calls per turn — eight is generous for most agents — bounds the worst turn's context growth, its concurrency against your dependencies, and the size of the mess a bad batch can make. It is the cheapest blast radius control in the tool layer.

Go and read one week of production transcripts for the widest batch your agent has ever emitted, and check it against the two properties in Step 2. Most teams find at least one batch containing two writes, and they find it in the logs rather than in an incident because nothing has yet retried at the wrong moment. Then add the reading/mutating boolean to your tool registry and make the dispatcher enforce it — an afternoon's work that turns a class of race conditions into a return value, and leaves the fan-out you actually wanted untouched.

Related: advanced tool orchestration patterns for the batch's more structured cousins, streaming tool calls for what the wire actually looks like while a batch is being emitted, and code as action for the approach that replaces the batch with a program the model can express dependencies in.