Deep-Dives / Multi-Agent Systems

Multi-Agent Systems

When multiple agents pay off, when one beats many, and the topologies, failure modes and debate patterns in between.

  1. When (and When Not) to Go Multi-Agent
    Price the coordination tax before you split: the three honest reasons to add an agent, and when one agent with tools wins.
  2. Multi-Agent Topologies
    Star, pipeline, hierarchy, mesh — their O(·) message cost and failure profiles, and how to pick the sparsest wiring that still works.
  3. Supervisor / Worker Orchestration
    The pattern that actually ships: plan, dispatch isolated workers, aggregate — and why the supervisor is the bottleneck.
  4. Debate, Voting & Ensembles
    Most of the gain is ensembling, not debate; without engineered diversity, debate collapses to the initial majority.
  5. Shared Memory & the Blackboard
    A blackboard replaces N² messages with one shared store — and inherits write contention, stale reads, and lost updates.
  6. Multi-Agent Failure Modes
    Error propagation, groupthink, deadlock/livelock, cost explosion — the system-level bugs single-agent tooling cannot see.
  7. Sub-Agent Patterns Compared
    LangGraph supervisor / hierarchical / collaborative vs OpenAI Agents SDK handoffs-vs-agents-as-tools vs deepagents — when each shape works.
  8. Credit Assignment: Which Agent Do You Change?
    A run-level score contains no gradient, so per-agent judges, ablation and counterfactual replay answer three different questions — local quality, marginal contribution, and specific causation — and reporting the cheapest as though it were the most specific is how a team spends a quarter tuning the worker that did exactly what it was told.
  9. What a Sub-Agent Actually Inherits
    Every framework gives you a switch for how much conversation a sub-agent inherits and none for how much authority, so the setting you tune costs tokens while the one you cannot see decides how bad a compromised run gets. Tools bind to the process, not the conversation: an isolated sub-agent is a fresh context window with identical reach, and fan-out buys concurrency at constant authority. Worse, isolation launders provenance — a brief written from a hostile page arrives in the child's instruction position looking trusted. Five inheritance decisions, one declaration per role.
  10. Isolated Runs That Find Each Other
    Isolation is enforced at the process boundary and specified at the information boundary, so any object two runs can write and read turns a fleet into one system with no supervisor and no aggregate budget. Convergent discovery makes the first occurrence and the hundredth minutes apart at fleet scale; five things then break, four of them silently, including the statistical independence your confidence interval assumed. Detect it in the join, namespace every writable object unguessably, and give the cohort an owner that can stop it.