0%
Applied AI Systems
Retrieval and Knowledge
Agents and Reasoning
Evaluation, Interpretability, and Safety
Production Research Engineering
Multi-Agent Patterns in Depth
The lesson on agent architectures introduced the idea that you can split a task across multiple agents, but it stopped at the level of a sketch. This lesson is the deeper playbook. When should you actually reach for a multi-agent design? Which pattern matches which problem? And where do these patterns break down in practice?
We start with the most common multi-agent structure: supervisor and worker. It is the default choice for roughly 70 percent of production multi-agent systems, and understanding its strengths and failure modes is the foundation for everything else in this lesson.
The Shape of the Pattern
A supervisor-and-worker system has one agent, the supervisor, whose job is to receive a task, decide which worker should handle each part of it, dispatch the subtasks, collect the results, and produce a final answer. The workers are specialized agents, each one narrowly focused on a single domain or tool set. The supervisor never does the domain work itself. It only routes and aggregates.
If you squint, it looks like a traditional software architecture with a load balancer in front of a pool of specialized services. That analogy is mostly correct, and it is why the pattern feels natural to engineers coming from distributed systems. The key difference is that the "load balancer" is an LLM making routing decisions in natural language, not a rule-based dispatcher.
Why Specialize Workers at All
You could put all the tools and all the instructions into a single agent and call it done. For small systems this often wins. So why split?
The reason is prompt bloat. Every tool description, every example, and every edge-case instruction sits in the system prompt. A single agent that can handle research, customer records, billing, and legal review ends up with a 5000-token system prompt that the model has to re-read on every turn. This hurts latency, cost, and accuracy. Models reliably start missing instructions once a system prompt exceeds a few thousand tokens, and the instructions most at risk are the rarely-used ones tucked in the middle.
Splitting into workers solves this. Each worker has a focused system prompt that describes only its job. The research worker knows about search APIs and citation format. The billing worker knows about the subscription schema and refund policies. Neither sees the other's instructions, and neither has to reason about irrelevant context. The supervisor is the only component that sees the full space of options, and its system prompt can be relatively small because it only needs to describe the workers, not what they do internally.
A useful mental model is to treat system prompt tokens as a finite budget the model spends on attention. Every irrelevant instruction taxes every decision. Splitting workers gives each one a clean tax-free prompt for its specific job.
Concrete Example: Multi-Domain Customer Support
Imagine a customer support system that handles account issues, billing disputes, technical troubleshooting, and refund requests. A single-agent design would need one mega-prompt covering all four domains plus all the tools for each. Users would experience slow responses (because every request processes the full prompt), inconsistent behavior across domains (because models struggle to apply the right policy to the right domain in one shot), and painful instruction drift as the system grows.
A supervisor-and-worker design has:
- A lightweight supervisor whose only job is to read the user message, classify the domain, route the message to the appropriate worker, and relay the worker's response back.
- An account worker with access to user record tools and identity-verification steps, plus a system prompt that covers only account policy.
- A billing worker with access to payment history tools, refund tools, and a prompt covering only billing edge cases.
- A technical worker with access to diagnostics tools, a knowledge base search tool, and a prompt covering only troubleshooting playbooks.
- A refunds worker with stricter guardrails and access to the refund approval tool.
Each worker is shorter, cheaper, and more accurate within its domain. The supervisor is cheap because it does not need to know the details of any single domain, only which worker owns which category.
Communication Protocols
How exactly does the supervisor talk to a worker? There are three common choices.
Structured handoff. The supervisor emits a JSON object with a worker name and a task description. The framework invokes that worker with a fresh context containing the original user request plus the supervisor's task description. The worker returns a JSON object with a result. This is the cleanest pattern and the one LangGraph favors.
Shared scratchpad. All agents share a single growing message list. The supervisor posts a message like "research worker, please look up the customer's recent orders." The worker sees the entire scratchpad and posts back a response. This is simpler but fragile at scale because every worker pays the token cost of every other worker's output. It is the default in early AutoGen examples.
Message-passing with topics. Each worker subscribes to a topic and the supervisor publishes to topics. This is what Microsoft Semantic Kernel's multi-agent framework uses. It is more complex but scales to systems with dozens of workers because no single conversation thread becomes the bottleneck.
For most systems, structured handoff is the right default. It keeps worker contexts small, makes debugging tractable (each worker invocation has a clean input and output), and maps naturally to tracing systems like LangSmith or Weights and Biases.
The Load-Balancing Problem
In traditional distributed systems, load balancing is a well-understood problem. A layer-4 balancer distributes connections by hash or round-robin. A layer-7 balancer inspects HTTP headers and routes by path. Either way, the routing decision is cheap and deterministic.
Multi-agent systems are different because the "routing decision" is an LLM call. It is expensive (one extra LLM invocation per request), non-deterministic (the supervisor might pick the wrong worker), and opaque (you cannot easily inspect the supervisor's internal reasoning). These three properties combine into a real engineering problem.
Cost. If your supervisor uses GPT-4 to decide between four workers, you have added one GPT-4 call to every request. For high-traffic systems this is significant. A common optimization is to use a smaller model for the supervisor (GPT-4o-mini or Claude Haiku) because classification is easier than generation. This can cut the supervisor cost by 10x without meaningful accuracy loss.
Non-determinism. The supervisor can misroute. A user asks about a billing issue but phrases it as "my account is broken," and the supervisor sends it to the account worker. The account worker gets confused, sends it back to the supervisor, and now you have two wasted LLM calls. Mitigation: give each worker an explicit escape hatch where it can signal "this is not my domain, please reroute." LangGraph supports this via conditional edges.
Opacity. When a system misroutes, you want to know why. The supervisor's decision is just a natural-language message. Good observability practice here is to have the supervisor emit a one-line justification with every handoff, then log that justification alongside the decision. Debugging routing failures becomes much easier when you can read why the supervisor picked a worker.
When to Use It and When Not
Use supervisor-and-worker when:
- Your task naturally splits into distinct domains with different tool sets or different policies.
- Your system prompt is pushing past 3000 tokens and growing.
- You have observability needs that benefit from clean per-worker traces.
- You anticipate the worker set growing over time, so a single mega-prompt would become unmaintainable.
Do not use supervisor-and-worker when:
- The task is simple enough that a single agent with three or four tools can handle it. Adding a supervisor just adds latency and cost.
- Subtasks have strong sequential dependencies where worker N needs the full state from worker N-1. The structured handoff pattern makes this awkward. Consider a single agent with good memory management instead.
- Latency is critical and you cannot afford the extra supervisor round trip. A 1-second request becomes a 2-3 second request when you add a supervisor step, and that matters for conversational UIs.
Frameworks That Implement This
LangGraph is the most popular multi-agent framework built around this pattern. Its create_supervisor helper gives you a supervisor node that routes between worker nodes along conditional edges. CrewAI models the same pattern with a "manager" agent and "crew members", which is identical semantically with different naming. AutoGen's GroupChat with a designated speaker selector is another variant. All three frameworks have converged on essentially the same structure because it maps cleanly to how teams of humans solve problems: one coordinator, several specialists, explicit handoffs.
When prototyping a supervisor-and-worker system, start with just two workers and a trivially simple supervisor. Expand the worker count only when you have a clear reason. Every new worker adds a row to the supervisor's routing matrix and makes the system harder to reason about.