When One Agent Is Not Enough

A single AI agent can handle a lot: answer questions, classify tickets, generate content, analyze data. But complex business workflows involve multiple specialized tasks that benefit from different models, different prompts, and different tools. Trying to make one agent do everything creates a bloated, unreliable system.

Multi-agent systems split complex workflows across specialized agents that collaborate through structured communication. A research agent gathers information. An analysis agent synthesizes findings. A writing agent produces the output. A review agent checks quality. Each agent is optimized for its specific task.

Think of multi-agent systems like a well-run team: specialists who are great at their jobs, communicating through clear handoffs, with a coordinator making sure nothing falls through the cracks.

Architecture Patterns

Sequential pipeline

The simplest multi-agent pattern: agents execute in order, each passing its output to the next. An SEO content pipeline might flow: keyword research agent → outline agent → writing agent → editing agent → SEO optimization agent.

Advantages: simple to build, easy to debug, deterministic execution order. Disadvantages: slow (no parallelism), single point of failure at each stage, no feedback loops.

Fan-out/fan-in

A coordinator agent distributes work to multiple specialized agents in parallel, then aggregates results. A competitive analysis system might fan out to: pricing agent, feature comparison agent, review sentiment agent, and market positioning agent — all running simultaneously against the same competitor.

Advantages: fast (parallel execution), naturally handles independent subtasks. Disadvantages: aggregation is hard (combining heterogeneous outputs), partial failures need graceful handling.

Supervisor pattern

A supervisor agent dynamically decides which worker agents to invoke based on the current state. Unlike pipelines, the execution path is not predetermined. The supervisor evaluates each agent's output and decides the next step: invoke another agent, request revision, or declare the task complete.

This is the most flexible pattern and the one I use most in production. It handles branching logic, retries, and dynamic workflows that cannot be predetermined at design time.

Debate pattern

Multiple agents with different perspectives argue about a decision, and a judge agent synthesizes the best answer. Useful for high-stakes decisions where you want to explore multiple viewpoints before committing.

Communication Design

How agents communicate determines system reliability. Unstructured message passing (just text) leads to information loss and misinterpretation. Structured communication with typed messages catches errors early.

Message contracts

Define explicit schemas for inter-agent messages. Each message type has required fields, optional fields, and validation rules. When an agent produces output that does not match the expected schema, the system catches it immediately rather than propagating garbage downstream.

  • Task assignment: What to do, input data, expected output format, deadline
  • Task result: Output data, confidence score, execution metadata (tokens used, latency, model)
  • Error report: What failed, error type, partial results if any, suggested retry strategy
  • Review request: Output to review, criteria, accept/reject/revise options

State management

Multi-agent workflows need shared state. A key-value store (Redis, Cloudflare KV) works for simple cases. For complex workflows with branching and parallel execution, use a proper workflow engine that tracks which agents have executed, what their outputs were, and what comes next.

Practical Implementation

Agent specialization

Each agent should have a focused prompt, appropriate model selection, and specific tools. Do not give every agent access to every tool — that increases cost and error surface.

  • Research agents: Haiku or Sonnet, web search tools, large context window
  • Analysis agents: Sonnet or Opus, calculation tools, structured output
  • Writing agents: Opus for quality, Sonnet for speed, no tools needed
  • Validation agents: Haiku for cost efficiency, structured rubrics, pass/fail output

Cost control

Multi-agent systems can get expensive fast. Three agents each making 5 LLM calls with Opus adds up. Cost control strategies:

  • Use the cheapest model that meets each agent's quality bar
  • Cache intermediate results to avoid redundant computation
  • Set per-agent token budgets with hard cutoffs
  • Monitor cost per workflow execution and alert on anomalies

Error handling

In a multi-agent system, failures cascade. One agent's bad output becomes the next agent's bad input. Build circuit breakers:

  • Validate every agent's output against its schema before passing downstream
  • Set retry limits per agent (typically 2-3 attempts)
  • Implement fallback paths: if agent A fails, can agent B produce an acceptable result?
  • Track failure rates per agent and alert when they exceed thresholds

Real-World Example: Content Production

I run a multi-agent content system that produces blog articles for 15 sites. The workflow:

  1. Topic agent (Haiku): analyzes the site's existing content, identifies gaps, proposes topics with SEO keywords
  2. Outline agent (Sonnet): creates a structured outline with H2/H3 headings, key points per section, and internal linking opportunities
  3. Writing agent (Opus): produces the full article following the outline, matching the site's voice and style
  4. SEO agent (Haiku): optimizes meta tags, validates JSON-LD, checks keyword density, adds internal links
  5. Deploy agent (no LLM): generates the HTML file, updates the blog index, updates the sitemap, deploys to Cloudflare Pages

The entire pipeline runs in under 3 minutes per article at approximately $0.15 per article. The supervisor agent monitors quality at each stage and can request revisions or abort if quality drops below threshold.

Common Mistakes

  • Over-engineering. Start with a sequential pipeline. Graduate to fan-out or supervisor only when you need the flexibility. Most workflows do not need debate patterns or complex coordination.
  • No observability. When a 5-agent workflow produces bad output, you need to see exactly what each agent produced. Log everything with correlation IDs.
  • Shared prompts. Each agent needs its own optimized prompt. Do not try to make one "universal agent" that switches roles based on instructions.
  • Ignoring latency. Sequential pipelines add latency with every agent. If your workflow has 5 agents each taking 3 seconds, the user waits 15 seconds. Parallelize where possible.

Getting Started

  1. Identify a workflow that currently uses one overloaded agent. Look for tasks where quality suffers because the agent is doing too many different things.
  2. Decompose into 2-3 specialized agents. Start small. Define clear input/output schemas for each.
  3. Build a sequential pipeline. Wire the agents together with schema validation between each step.
  4. Add a supervisor once the pipeline works. Let it handle retries and quality checks.
  5. Monitor cost and latency per agent. Optimize the expensive ones first.

Related Articles

WorkflowOrchestration

AI Workflow Orchestration

Building reliable multi-step AI workflows with state management and error recovery.

Autonomous AIScheduled Agents

Building an Autonomous AI Operator

Scheduled AI agent loops that run business operations without human intervention.

Cost OptimizationLLM

Your AI Bill Is 10x What It Should Be

Model routing, caching, and optimization that cut AI costs by 70-90%.