Why Single-Agent AI Breaks Down
Most AI automations start as a single LLM call: take input, generate output, done. This works for simple tasks — summarize an email, classify a support ticket, extract data from a document. But real business workflows have branching logic, human-in-the-loop checkpoints, external API calls, and state that needs to persist across hours or days.
The moment you chain three or more LLM calls together, you hit reliability problems. Each call has a failure rate. Token limits truncate context. Intermediate results need validation before the next step can proceed. A single-agent approach either becomes a fragile script full of try/except blocks or an unreliable mess that silently produces garbage.
The shift from "AI tool" to "AI system" is the shift from a single prompt to an orchestrated workflow. Most teams get stuck here because they try to solve orchestration problems with better prompts instead of better architecture.
The Workflow Orchestration Pattern
Instead of one agent doing everything, you decompose the workflow into discrete steps, each with a clear input, output, and success criteria. An orchestrator manages the flow, handles failures, and maintains state.
Step decomposition
Take any business process you want to automate and break it into atomic steps. Each step should:
- Have a single, testable responsibility
- Accept structured input and produce structured output
- Be independently retryable without side effects
- Include validation criteria that determine pass/fail
Example: an automated proposal generator for a consulting firm. The naive approach is one prompt: "Write a proposal for [client]." The orchestrated approach:
- Research step: Pull client info from CRM, recent communications, industry data
- Scope step: Generate a scope of work based on the research + service catalog
- Pricing step: Calculate pricing using rate cards, historical margins, client tier
- Draft step: Generate the proposal document combining scope + pricing + boilerplate
- Review step: AI quality check for consistency, completeness, pricing accuracy
- Human approval: Route to the right partner for sign-off
State management
Every workflow needs persistent state. I use a simple pattern: each workflow instance gets a JSON document that accumulates results as steps complete. The orchestrator reads state, determines the next step, executes it, and writes the result back.
class WorkflowState:
def __init__(self, workflow_id, workflow_type):
self.id = workflow_id
self.type = workflow_type
self.steps = {}
self.current_step = None
self.status = "running" # running | waiting | completed | failed
self.created_at = datetime.utcnow().isoformat()
def complete_step(self, step_name, result):
self.steps[step_name] = {
"status": "completed",
"result": result,
"completed_at": datetime.utcnow().isoformat()
}
def fail_step(self, step_name, error, retries=0):
self.steps[step_name] = {
"status": "failed",
"error": str(error),
"retries": retries,
"failed_at": datetime.utcnow().isoformat()
}
Error recovery
Every step can fail. The orchestrator needs a recovery strategy for each failure mode:
- Transient failures (API timeout, rate limit) — retry with exponential backoff, up to 3 attempts
- Validation failures (output does not meet criteria) — retry with modified prompt or additional context
- Hard failures (missing data, permission denied) — pause workflow, notify human, resume when resolved
- Quality failures (output is technically valid but poor quality) — escalate to human review with the AI's self-assessment
Real Workflow: Automated Client Onboarding
Here is a production workflow I built for a service business:
- Intake: New client fills out a form. The AI extracts structured data (company size, industry, needs, budget range, timeline).
- Qualification: AI scores the lead against ideal customer profile. Scores below threshold get a polite automated response. Above threshold continues.
- Research: AI researches the company (website, LinkedIn, news, competitors). Produces a 1-page brief.
- Proposal draft: AI generates a customized proposal using the brief + service catalog + pricing rules.
- Internal review: AI checks the proposal for pricing errors, scope gaps, and compliance issues. Flags anything unusual.
- Human sign-off: Partner reviews the flagged items and the proposal. Approves, requests changes, or rejects.
- Delivery: Approved proposal is formatted as PDF, personalized email is drafted, and sent via the CRM.
- Follow-up: If no response in 3 days, AI drafts a follow-up email referencing specific points from the proposal.
Total time from form submission to proposal delivery: under 2 hours (including the human review step). Previous manual process: 2-3 days. The AI handles steps 1-5 and 7-8 autonomously. The human only touches step 6.
Choosing the Right Orchestration Tool
You do not need a framework to orchestrate AI workflows. A Python script with a state machine and a task queue is often sufficient for simple workflows. But as complexity grows, consider:
- Simple (3-5 steps, linear): Python script + SQLite state + cron. No framework needed.
- Medium (5-15 steps, branching): Temporal.io or a custom state machine. Temporal handles retries, timeouts, and human-in-the-loop natively.
- Complex (15+ steps, parallel branches, long-running): Temporal.io with activity workers or a custom DAG executor. At this scale, you also need observability (logging every step's input/output for debugging).
Common Mistakes
- Over-engineering the orchestrator before building the steps. Get each step working independently first. Chain them manually. Only then build the orchestrator.
- Passing raw LLM output between steps. Always validate and structure the output of each step before passing it to the next. Garbage propagates.
- Skipping the human-in-the-loop checkpoint. Every workflow that produces customer-facing output needs at least one human review point until you have months of data proving the AI's accuracy.
- Not logging intermediate state. When a 6-step workflow produces a bad result, you need to inspect each step's input and output to find where it went wrong. Log everything.
- Building for the happy path only. The orchestrator's value is almost entirely in how it handles failures. If your workflow only works when every API call succeeds and every LLM output is perfect, it is not production-ready.
Getting Started
- Pick one workflow that currently takes a human 2+ hours and involves at least 3 distinct steps.
- Document each step with its input, output, and success criteria. Be specific about what "good enough" looks like.
- Build each step as an independent function that you can test in isolation. Use real data, not synthetic examples.
- Chain them with a simple script that runs steps sequentially, validates output, and logs everything.
- Add error handling after the happy path works. Start with retries, then add human escalation for unrecoverable failures.
Related Articles
Building an Autonomous AI Operator
Scheduled AI agent loops that run business operations without human intervention.
AI Customer Support That Actually Resolves Tickets
Step-by-step guide to building AI support that resolves 80% of inquiries automatically.
Your AI Bill Is 10x What It Should Be
Model routing, caching, and optimization techniques that cut AI costs by 70-90%.