Skip to content

Group: Agent Systems | Exit of the group above (Action): you can execute a single tool call safely | Exit of this page: you can judge when multi-agent is worth it, write the four-element delegation contract, delegate to specialists via a supervisor topology, and degrade gracefully on routing failure instead of crashing Prerequisites: Agent Runtime, Workflow Patterns | Next: A2A (protocols only across boundaries), Observability, Cost and Performance

1. Overview ​

BLUF: multi-agent is not "a stronger single agent" — it is an architecture that trades coordination for capacity. Independent context windows enable parallel exploration and compression (subagents distill vast raw material into conclusions and return them), buying coverage a single agent cannot reach; the price is roughly 15× the tokens of a chat, coordination complexity, and error propagation. The test has three clauses, and all must hold: the task is valuable enough to pay for it, the sub-directions are naturally parallel, and the information exceeds a single context. Missing any, go back to a workflow or a single agent.

Scope up front: this page covers only orchestration of subagents spawned inside one trust domain (same host / process); delegating tasks to agents across processes, organizations, or trust domains is not expanded here — go through A2A first, picking the connection direction on the Protocol Map (criteria in "Key boundary" below).

Mental model: supervisor + specialists + artifacts ​

Three components: the supervisor (decomposes tasks, writes delegation briefs, synthesizes results), specialists (subagents that each own an isolated context window, toolset, and prompt), and artifacts (large outputs land in storage and only references travel back — avoiding information loss from multi-hop retelling).

When to use / when not to ​

  • Use (Anthropic's production framing): high-value tasks + heavy parallelism + information beyond one context + many complex tools to interface with; the archetype is open-ended research (breadth-first queries).
  • Skip: tasks where all agents must share one context, or where inter-agent dependencies are dense — most coding tasks lack true parallelism and fit poorly; real-time coordination and delegation between agents is not yet a strength. Enumerable steps belong in a workflow; single-context problems belong to a single agent.

Decision table: ladder placement ​

OptionDirectionControlStateTrust domainMinimum complexity
Single agent (Agent Runtime)Read-write loopModel + stopping conditionsOne contextWithin hostDefault starting point
Same-domain multi-agent (this page)Read-write loop × NSupervisor delegates and routesPer-agent contexts + synthesis pointSame host/processParallelism/specialization beats coordination cost
Cross-boundary collaboration (Protocol Map)Read-write across boundariesProtocol negotiation (A2A etc.)Tasks/messages/artifactsCross-process/org/trust domainNeeded only when a second trust domain appears

Key boundary: subagents spawned inside one process do not need a protocol — function calls and structured briefs suffice; only collaboration that crosses processes, organizations, or trust domains enters the protocol branch such as A2A.

Gains vs costs ​

DimensionGainCost
ContextParallel exploration in separate windows; compressed returnsSynthesis-point bottleneck; retelling loss (game of telephone)
SpecializationDistinct tools/prompts/trajectories reduce path dependencyTools and prompts to maintain ×N
Parallelism3–5 subagents at once; time cut by up to 90%Total tokens ≈15× chat; needs budget guardrails
ReliabilityA failed subagent can be rescued by the supervisorErrors propagate across agents; emergent behavior is hard to predict
DebuggingDelegation traces shard naturallyThe global causal chain must be assembled across shards

Orchestration topologies ​

TopologyStructureFitsFailure mode
Supervisor (orchestrator-workers)A central agent decomposes and delegates, then synthesizesOpen tasks with non-predefinable subtasksSupervisor becomes the information bottleneck; sequential waiting
RouterClassify, then dispatch to a specialized agent; no synthesisTriage with clear input classesSilent failure on misclassification
PipelineFixed order, intermediate products passed alongStage-explicit processing chainsRetelling loss amplifies stage by stage
Debate/verify (generator-verifier)Generate-and-evaluate pairs in a loopQuality-critical outputs with clear criteriaOscillation without convergence; needs a max-iteration cap

Anthropic's five patterns (generator-verifier / orchestrator-subagent / agent teams / message bus / shared state) expand these four; evolution criteria are in the resource library.

Historical milestones: this page consolidates the old "Multi-Agent Coordination Patterns" page (a repository archive of a Claude blog post, 2026-04-10); the five patterns remain in the resource library while the body reorganizes around four topologies and adds the delegation contract and degraded-fallback semantics. Earlier timelines unverified; not fabricated.

2. Usage ​

Minimal hands-on: a supervisor-pattern mock — two specialist agents receive structured briefs (the four TaskBrief elements) and return results, demonstrating normal routing, degraded fallback on routing failure, and contract rejection. Specialists are deterministic mocks, not LLMs: what is being verified is the orchestration layer's contract and degradation, not model capability. Zero API keys, zero dependencies.

Environment: Node ≥ 22.18. Save as multi-agent.ts, run node multi-agent.ts.

ts
// multi-agent.ts — a supervisor-pattern mock: delegation contract + specialist routing + degraded fallback.
// Zero dependencies, zero API keys (specialists are deterministic mocks, not LLM calls).
// Runs natively on Node >= 22.18: node multi-agent.ts
import { setTimeout as sleep } from "node:timers/promises";

// ---------- Delegation contract: four required fields (objective / outputFormat / boundaries / taskType) ----------
interface TaskBrief {
  taskType: "search" | "summarize";
  objective: string;     // what question to answer
  outputFormat: string;  // what structure to return
  boundaries: string;    // what to do and what NOT to do
}

interface SpecialistResult {
  specialist: string;
  payload: unknown;
}

// ---------- Specialist agents: each with its own isolated context (mocked via closures) ----------
const specialists: Record<TaskBrief["taskType"], {
  name: string;
  handle: (brief: TaskBrief) => Promise<SpecialistResult>;
}> = {
  search: {
    name: "search-specialist",
    handle: async (brief) => {
      await sleep(10);
      // real system: this is a subagent loop with its own context window
      return { specialist: "search-specialist", payload: { query: brief.objective, hits: ["doc-a", "doc-b"] } };
    },
  },
  summarize: {
    name: "summarize-specialist",
    handle: async (brief) => {
      await sleep(10);
      return { specialist: "summarize-specialist", payload: { summary: "2 findings: cost is dominated by tool calls; retries need idempotency keys." } };
    },
  },
};

// ---------- Supervisor: validate the delegation contract -> route -> degrade on routing failure ----------
class Supervisor {
  private trace: string[] = [];

  async delegate(brief: TaskBrief): Promise<{ mode: "delegated" | "degraded" | "rejected"; result?: unknown; reason?: string }> {
    // 1) Delegation-contract check: a brief without an objective is rejected outright
    //    (prevents subagent spin-up with no steering, or duplicated work)
    if (!brief.objective || brief.objective.trim().length === 0) {
      this.trace.push("rejected: brief missing objective");
      return { mode: "rejected", reason: "brief missing objective — subagent cannot be steered" };
    }
    const specialist = specialists[brief.taskType];
    // 2) Degraded fallback on routing failure: the supervisor answers itself instead of crashing
    if (!specialist) {
      this.trace.push(`degraded: no specialist for taskType=${brief.taskType}`);
      return {
        mode: "degraded",
        result: { answer: `supervisor fallback for "${brief.objective}" (no specialist: ${brief.taskType})` },
      };
    }
    // 3) Normal delegation: structured task in, structured result out
    const result = await specialist.handle(brief);
    this.trace.push(`delegated: ${brief.taskType} -> ${result.specialist}`);
    return { mode: "delegated", result: result.payload };
  }

  getTrace() { return this.trace; }
}

async function main() {
  const supervisor = new Supervisor();

  console.log("== case 1: search task -> routed to search-specialist ==");
  const r1 = await supervisor.delegate({
    taskType: "search", objective: "find docs about agent cost",
    outputFormat: "list of doc ids", boundaries: "only internal wiki, max 5 queries",
  });
  console.log(JSON.stringify(r1));

  console.log("== case 2: summarize task -> routed to summarize-specialist ==");
  const r2 = await supervisor.delegate({
    taskType: "summarize", objective: "summarize search findings for the cost report",
    outputFormat: "3 sentences", boundaries: "no new searches",
  });
  console.log(JSON.stringify(r2));

  console.log("== case 3: unknown taskType=translate -> routing failure degrades gracefully ==");
  const r3 = await supervisor.delegate({
    taskType: "translate" as any, objective: "translate the report to English",
    outputFormat: "translated text", boundaries: "keep terminology",
  });
  console.log(JSON.stringify(r3));

  console.log("== case 4: negative — a brief missing its objective is rejected by the contract ==");
  const r4 = await supervisor.delegate({
    taskType: "search", objective: "", outputFormat: "list", boundaries: "wiki only",
  });
  console.log(JSON.stringify(r4));

  console.log("== supervisor trace (failure localization: who handled what) ==");
  for (const line of supervisor.getTrace()) console.log(`  ${line}`);
}

main();

Normal output (deterministic):

text
== case 1: search task -> routed to search-specialist ==
{"mode":"delegated","result":{"query":"find docs about agent cost","hits":["doc-a","doc-b"]}}
== case 2: summarize task -> routed to summarize-specialist ==
{"mode":"delegated","result":{"summary":"2 findings: cost is dominated by tool calls; retries need idempotency keys."}}
== case 3: unknown taskType=translate -> routing failure degrades gracefully ==
{"mode":"degraded","result":{"answer":"supervisor fallback for \"translate the report to English\" (no specialist: translate)"}}
== case 4: negative — a brief missing its objective is rejected by the contract ==
{"mode":"rejected","reason":"brief missing objective — subagent cannot be steered"}
== supervisor trace (failure localization: who handled what) ==
  delegated: search -> search-specialist
  delegated: summarize -> summarize-specialist
  degraded: no specialist for taskType=translate
  rejected: brief missing objective

What to look at: case 3's mode:"degraded" — the routing failure did not throw or crash; the supervisor answered itself and left a trace. Case 4's mode:"rejected" — the delegation contract turned an unsteerable brief away before spawning anything.

Acceptance command:

bash
node multi-agent.ts | grep -c '"mode"'   # expected: 4 (delegated ×2 / degraded / rejected)

Cleanup: purely in-memory; no side effects.

Scenario walkthrough ​

ScenarioInputActionOutputFitsDoes not fit
Open researchOne broad questionSupervisor splits 3–5 sub-directions, delegates in parallelSynthesis + citationsBreadth-first, beyond one contextQueries with a fixed answer chain (a single agent is cheaper)
Multi-dimension reviewOne documentOne specialist per dimension (security/perf/style)Per-dimension findingsIndependent, parallelizable dimensionsTightly coupled dimensions (they overturn each other)
Support triageUser requestRouter classifies and dispatchesSpecialized path handles itClearly classifiable inputsAmbiguous classes (silent misclassification failures)

3. Principles ​

Why multi-agent works: capacity and compression ​

Anthropic's production data analysis: on the BrowseComp benchmark, token usage alone explains 80% of performance variance (with tool-call count and model choice, 95%). The essence of the multi-agent architecture is scaling token usage for tasks beyond a single agent's limits: each subagent explores in its own context window and returns the most important tokens after compression. In internal evals, an Opus 4 lead + Sonnet 4 subagents beat single-agent Opus 4 by 90.2% (breadth-first research queries).

The delegation contract: teaching the supervisor to delegate ​

Subagent output quality is capped by the delegation brief. Anthropic's minimal four-element contract (a missing element is case 4's rejected):

ElementQuestion it answersConsequence if missing
objectiveWhat to answerThe subagent cannot be steered; it spins or duplicates others' work
outputFormatWhat structure to returnThe synthesis side fails to parse; retelling loss amplifies
Tool & source guidanceWhat to use, what not toSearching the web for facts that only exist in Slack
boundariesDo what, don't do whatSeveral subagents collide on the same work

With matching scaling rules (written into the supervisor prompt): simple fact-finding, 1 agent with 3–10 tool calls; direct comparisons, 2–4 subagents with 10–15 calls each; complex research, 10+ with clear division. Early systems without scaling rules once spawned 50 subagents for a simple query.

State: shared or isolated ​

ModeMechanismFitsRisk
Isolated + message returnSubagents communicate only via structured results to the supervisorDefault: exploratory, compressible tasksSynthesis-point bottleneck
Shared storeAgents read/write the same files/DB/knowledge baseCollaborative construction (findings influence each other)Duplicated work; reactive loops (A writes → B responds → A responds again, burning tokens indefinitely) — needs explicit termination conditions
Artifact bypassLarge outputs go straight to the filesystem; only references returnLong reports, code, datasetsReference decay needs governance

Failure localization ​

The debugging unit of a multi-agent system is the delegation: every delegate records a brief digest, routing outcome, and subagent terminal state. Behavior is emergent — a small supervisor-prompt change can unpredictably change subagent behavior — so evaluation targets the end state, not the step-by-step path: assert the final state is correct rather than the path matching a preset. Cross-shard causal chains are assembled from traces (into the Production group's Observability).

Spec requirements vs local test ​

Official statement (Anthropic's multi-agent research-system retrospective)Local fixture counterpart
Orchestrator-worker pattern: the lead agent plans and spawns parallel subagents (retrievedAt 2026-09-01)Supervisor.delegate routes to specialists
Delegation briefs need objective, output format, tool/source guidance, task boundariesThe four TaskBrief fields; missing objective → rejected
Simple queries once spawned 50 subagents → scaling rules embeddedThis page lists scaling rules in Principles (not implemented in the fixture; it is a prompt-layer concern)
Token usage: agents ≈ 4× chat, multi-agent ≈ 15× chatThe fixture's mock specialists cost zero tokens — the decision tables keep the constraint
Large outputs written to the filesystem with references returned, reducing retelling lossThe artifacts bypass in the mental-model diagram

4. Development ​

Integration ​

  1. The supervisor is itself, in the sense of Tool Execution Engineering, a tool below the approval tier: the spawn-subagent action passes an allowlist and a budget gate.
  2. Each specialist is an independent agent loop (see Agent Runtime): its own context, toolset, and prompt; no shared memory.
  3. When a delegation product crosses a threshold (say 2K tokens), switch to the artifact bypass: write a file, return a path reference.
  4. Budget guardrails: per-run subagent cap, per-subagent tool-call cap, total token cap — any one tripping stops further spawns and synthesizes what exists.

Testing ​

  • Three routing branches: deterministic assertions for delegated / degraded / rejected (the fixture is the template).
  • Contract negatives: a brief missing each element must be rejected.
  • End-state evaluation: assert the final product end-to-end, not the intermediate path (multi-agent paths are non-deterministic).

Rollback ​

Rolling back the orchestration layer = reverting the supervisor prompt and routing table. Beware emergence: a small change can largely alter subagent behavior; after rollback, re-run the end-state eval set instead of eyeballing one example.

Symptom → Evidence → Fix → Done ​

Symptom: two subagents return nearly identical content, and a third direction goes uncovered. Evidence: the two briefs' objectives overlap heavily and neither has boundaries; traces show repeated search terms. Fix: complete the delegation contract — split objectives to be mutually exclusive, write "do not do X" explicitly in boundaries; add a "self-check coverage after splitting" step to the supervisor prompt. Done: the brief set for one task has pairwise non-overlapping objectives; duplicate retrieval disappears from traces.

Symptom → Evidence → Fix → Done ​

Symptom: even simple questions spawn a dozen subagents; the bill explodes. Evidence: no scaling rules; trace shows subagent count uncorrelated with question complexity. Fix: write the "1 / 3–10, 2–4 / 10–15, 10+" scaling ladder into the supervisor prompt; add a hard per-run spawn cap. Done: subagent count for simple queries stabilizes at 1; the cap gate has test coverage.

Symptom → Evidence → Fix → Done ​

Symptom: the supervisor's synthesis loses detail, even distorts subagent findings. Evidence: subagent raw outputs are long and all transit through the supervisor's context (retelling loss). Fix: switch large products to the artifact bypass — subagents write directly to storage and return a reference plus a digest; the supervisor reads on demand. Done: key facts in the final product trace back to artifact originals; supervisor context usage drops noticeably.

Symptom → Evidence → Fix → Done ​

Symptom: after a multi-agent failure, nobody can locate which link introduced the error. Evidence: only final-answer logs exist; no per-delegation traces. Fix: record per delegate (brief digest, routing decision, subagent terminal state, artifact reference); switch evaluation to end-state assertions plus trace sampling. Done: any failure can be attributed to a specific delegation in the trace; the end-state eval set is green.

Anti-patterns ​

  • Multi-agent for scale's sake: tasks without parallelism (most coding) harvest only coordination overhead.
  • One-line delegation: "research the semiconductor shortage"-style briefs breed duplicated work — all four contract elements, no exceptions.
  • Unguarded spawning: any "spawn another on demand" path needs a hard cap.
  • Shared state without termination conditions: reactive loops burn tokens until the budget hits zero.
  • Multi-agent as a fix for single-agent prompt problems: fix the single agent's tool descriptions and prompts before splitting — splitting copies the problem rather than solving it.

5. Resource Library ​

Four-level reading route:

  • Beginner: finish this page → recite "the three-clause test + the four delegation elements"; run the fixture's three-branch output.
  • Builder: Agent Runtime (what a specialist is) + integrate this page's contract code into a supervisor.
  • Operator: the production sections of the multi-agent research retrospective (checkpoints, rainbow deployment, end-state evaluation); token budgeting in Cost and Performance.
  • Researcher: evolution criteria across the five coordination patterns; formal analysis of shared state and reactive loops.

Resource table ​

NameEvidence tiercanonical URLUseSupported claimNext
How we built our multi-agent research system (Anthropic)L1 (maintainer)https://www.anthropic.com/engineering/built-multi-agent-research-systemArchitecture, delegation contract, scaling rules, production reliability"Token usage explains 80% of variance; agents ≈4×, multi-agent ≈15× chat; parallelism cut time by up to 90%" (retrievedAt 2026-09-01)Read the production-reliability section closely
Multi-agent coordination patterns (Claude blog)L1 (maintainer)https://claude.com/blog/multi-agent-coordination-patternsFive patterns and pairwise evolution criteriaFive coordination patterns (verified via this repo's old-page archive 2026-04-10; not re-verified this round)Map onto this page's four topologies
Building multi-agent systems: when and howL1 (maintainer)https://claude.com/blog/building-multi-agent-systems-when-and-how-to-use-themPre-investment judgment"When multi-agent is worth it" (cited via the old page; not re-verified this round)Cross-check the three-clause test
Building Effective Agents (Anthropic)L1 (maintainer)https://www.anthropic.com/engineering/building-effective-agentsPositioning of the orchestrator-workers patternThe boundary where an orchestrator decomposes subtasks dynamically (retrievedAt 2026-09-01)Workflow Patterns
Learn LLM (sibling site)siblinghttps://llm.zenheart.site/Model-side roots of multi-agent behaviorModel mechanics belong to Learn LLM (retrievedAt 2026-09-01)Stop points above

Active falsification and open questions ​

  • Falsification entry: if you run a dependency-dense task on multi-agent both well and cheaply — share the task shape and bill comparison, and the three-clause test of this page needs revision.
  • Open: coordination semantics for asynchronous multi-agent (subagents communicating while running in parallel) are still evolving per Anthropic itself; this page covers only the synchronous supervisor pattern.
  • Open: how identity, authorization, and settlement for cross-organization multi-agent map onto the A2A task model — to be expanded in the Protocol Map branch.

Where learn-ai stops / where to go next ​

  • What a specialist is — the single-agent loop and state memory: Agent Runtime.
  • Cross-process/org/trust-domain agent collaboration: A2A (read the Protocol Map first to pick by connection direction).
  • Delegation-level traces and end-state evaluation: Observability, evals (Production group).
  • Accounting for the 15× tokens: Cost and Performance.

Built for frontend engineers · Powered by VitePress