Interest in multi-agent systems has exploded, and frameworks make it easy to wire up a planner, a researcher, a coder and a reviewer that talk to each other. It looks impressive. It is also frequently the wrong starting point.

Every extra agent adds latency, token cost, hand-off errors and another thing to debug. Start simple and add agents only when you can name the problem they solve.

The ladder of complexity

Climb one rung at a time.

  1. A single model call with a good prompt. Many "agent" problems are really this.
  2. A workflow: fixed steps in code, with a model call at some of them. Predictable and easy to test.
  3. A single agent with tools, choosing its own steps in a loop.
  4. Multiple agents, each with its own context and tools, coordinated by an orchestrator.

When multiple agents genuinely help

  • Context isolation. A sub-task produces lots of noise (web research, log analysis) that would pollute the main agent's context. A sub-agent digests it and returns a summary.
  • Parallel work. Independent sub-tasks can run at the same time, such as researching five competitors.
  • Different permissions. A reader agent can access untrusted content with no write tools; a separate agent with write access only sees the reader's vetted summary. This is a useful defence against prompt injection.
  • Specialised tool sets. Each agent gets a small, focused toolset instead of one agent choosing from fifty.

When they hurt

  • Agents need to share lots of state, so every hand-off loses detail.
  • The task is sequential and each step depends tightly on the previous one.
  • You cannot yet evaluate the single-agent version. Adding agents multiplies what you cannot measure.

Design the hand-offs

Most multi-agent failures happen at the boundaries.

  • Define a structured result for each sub-agent: what it found, confidence, sources, open questions.
  • Give sub-agents a clear stopping condition and a budget of steps or tokens.
  • Keep the orchestrator responsible for the final answer, so there is one place to check quality.
json
{
  "task": "Summarise breaking changes in library X v5",
  "result": ["Removed default export", "Node 18 no longer supported"],
  "sources": ["CHANGELOG.md#v5.0.0"],
  "confidence": "high",
  "open_questions": []
}

Measure before and after

Compare the multi-agent design against the single-agent baseline on the same eval set: success rate, cost per successful task and latency. If the gain is small, keep the simpler system.

Key takeaways

  • Start with a single call or a fixed workflow, then climb only when needed.
  • Multi-agent designs shine for context isolation, parallelism and permission separation.
  • Structured hand-offs and step budgets prevent most coordination failures.
  • Prove the gain with evals before accepting the extra cost.