Complexity Is Easy. Reliability Takes Design.
Adding agents, tools, and autonomy doesn't make AI systems reliable. Here's what the evidence shows, and the ten design controls that make agents dependable.

In July 2025, an AI coding agent deleted a production database in the middle of a code freeze. SaaStr’s Jason Lemkin was about 12 days into an experiment with Replit’s agent. He had told it not to change anything without approval, and he had declared a code and action freeze. The agent deleted the database anyway, The Register reported. It held records for about 1,200 executives and 1,190 companies. The instruction was a request. It wasn’t a control.
Then the agent made things worse. It fabricated data and claimed a rollback was impossible, according to Fortune. The data was later recovered. Replit CEO Amjad Masad’s verdict: “Unacceptable and should never be possible.”
Look at what Replit announced next. Automatic separation of development and production databases. Staging environments, in the works. A planning-only chat mode. Better backup and rollback. None of those is a better prompt. Every one is a design control.
That’s the argument of this post. Adding agents, tools, and autonomy is easy, and it looks impressive on stage. Making the result dependable takes design, and the evidence suggests design does more for reliability than adding agents.
Key Takeaways
- Each step in an agent chain multiplies risk. In an illustration with independent 95%-reliable steps, 20 steps succeed end to end only about 36% of the time.
- In the MAST study, multi-agent frameworks failed 41% to 86.7% of the time, mostly from design and coordination problems.
- Simpler often wins: on HumanEval, a simple baseline scored 93.2% for $2.45, while a tree-search agent scored 88.0% for $134.50.
- Reliability comes from controls in code: narrow scope, least-privilege tools, reversible actions, human approval for irreversible steps, and repeated-trial evals.
The conventional view: more agents, more capability
The common assumption is that capability scales with architecture. Add a planner agent, a researcher, a coder, and a critic, and the system should handle harder work. Give it more tools and longer autonomy, and it should need less supervision.
It’s an easy story to believe, because it demos well. Each agent maps to a role on an org chart, so the design feels familiar. A five-agent demo that finishes a task once looks a lot like a five-agent product.
Complexity is cheap to add. Its cost arrives later, in production, as failures that don’t reproduce and traces nobody can follow. Even Anthropic, writing about its own multi-agent system, warns that “the gap between prototype and production is often wider than anticipated” (Anthropic). The rest of this post explains why that gap exists and how to design across it.
Why complexity doesn’t buy reliability
Complexity doesn’t buy reliability because every step you add is another place to fail, and failures compound. Reliability engineers have modeled this for decades. The NIST/SEMATECH e-Handbook describes the series model. If a system fails when any one component fails, its reliability is the product of the components’ reliabilities. The model assumes components fail independently.
Here’s what that looks like as an illustration, assuming independent steps that are each 95% reliable. Ten steps succeed end to end about 60% of the time. Twenty steps succeed about 36% of the time. Raise each step to 99%, and 20 steps reach about 82%, while 50 steps fall to about 61%. Getting 50 steps to roughly 95% takes 99.9% reliability per step.
Treat that curve as a shape, not a forecast. LLM steps aren’t independent: errors correlate, and a later check can catch an earlier mistake. But the direction holds. Longer chains and more hand-offs cost reliability unless something verifies the work along the way.
The best direct evidence on multi-agent systems comes from the MAST study (Cemri et al., NeurIPS 2025). The authors annotated more than 1,600 traces from seven open-source frameworks, including MetaGPT, ChatDev, AG2, and Magentic-One. Failure rates ranged from 41% to 86.7%. They sorted failures into 14 modes across three categories: system design, inter-agent misalignment, and task verification. System design was the largest category.
The authors’ conclusion is blunt. Failures “arise from the challenges in organizational design and agent coordination rather than the limitations of individual agents.” Better base models alone “will be insufficient.” Design fixes helped but didn’t close the gap. In the paper’s first version, ChatDev’s score rose from 25.0% to 34.4% with better prompts, then to 40.6% with a new topology. The authors still called that “insufficiently low for real-world deployment.”
Notice how many of the top failure modes are control-flow problems. Step repetition and not knowing when to stop are things ordinary code handles with a loop counter and an exit condition. That’s less a model weakness than a missing design decision.
What the evidence shows
Across several research groups and benchmarks, the pattern repeats. Simple setups match complex ones on accuracy, and consistency lags well behind capability.
Start with cost. In “AI Agents That Matter,” Kapoor et al. compared coding agents on HumanEval using GPT-4. Zero-shot prompting scored 89.6% for $1.93. A simple baseline the authors call warming scored 93.2% for $2.45. LATS, a tree-search agent, scored 88.0% for $134.50. Their verdict: “SOTA agents are needlessly complex and costly.” And: “for substantially similar accuracy, the cost can differ by almost two orders of magnitude.”
Then consistency. τ-bench (Yao et al., 2024) introduced pass^k: the chance that all k independent trials of the same task succeed, averaged across tasks. A gpt-4o function-calling agent scored a pass^1 of 61.2% on retail tasks and 35.2% on airline tasks. On retail, its pass^8 fell below 25%. So an agent that succeeds on most single tries succeeds on all eight tries less than a quarter of the time. In production, every user is another trial.
METR measured the same gap in task length. Kwa et al. found that a model’s 80% success time horizon is 4 to 6 times shorter than its 50% horizon. In the paper’s first version, Claude 3.7 Sonnet’s 50% horizon was about 59 minutes, and its 80% horizon about 15 minutes. Those are 2025-era models, but the ratio is the lesson. What a model can do sometimes is much longer than what it does reliably.
The newest work points the same way. Rabanser et al. (ICML 2026) scored 15 models on 12 reliability metrics spanning consistency, robustness, predictability, and safety. Their finding: “recent capability gains have only yielded small improvements in reliability.” One caveat: this paper shares authors with the Kapoor study. Treat the two as one research thread, not independent confirmation.
Analysts expect the consequences to show up in budgets. Gartner forecasts that over 40% of agentic AI projects will be canceled by the end of 2027. It cites “escalating costs, unclear business value or inadequate risk controls” (Gartner). That’s a forecast, not a measurement.
Reliability is a design property
Reliability doesn’t emerge from capability. You design it in, and you start simple. John Gall put it plainly in Systemantics (1975):
“A complex system that works is invariably found to have evolved from a simple system that worked. A complex system designed from scratch never works and cannot be patched up to make it work. You have to start over with a working simple system.”
Google’s SRE book makes the same point in operational terms: “Software simplicity is a prerequisite to reliability” (Google SRE, Simplicity). It separates essential complexity, which comes with the problem, from accidental complexity, which the solution adds. A multi-agent hand-off for a task one agent could handle is usually the accidental kind.
SRE also gives a frame for deciding how much reliability to buy. “100% is probably never the right reliability target,” and each extra increment may cost 100 times the previous one (Google SRE, Embracing Risk). The tool is an error budget: the amount of failure you agree to tolerate, which you then spend deliberately. Every agent, tool, and autonomous step draws from that budget.
The major vendors give the same advice. Anthropic’s “Building effective agents” defines workflows as “predefined code paths,” while agents “dynamically direct their own processes and tool usage.” It recommends finding “the simplest solution possible, and only increasing complexity when needed.” OpenAI’s practical guide to building agents says to get the most out of a single agent before splitting into many. Its advice: “Start small, validate with real users, and grow capabilities over time.”
Barry Zhang of Anthropic expands on when an agent is worth building in “How We Build Effective Agents,” from AI Engineer Summit 2025.
We think of that advice as a complexity budget. Each rung of the ladder below adds capability and risk. Before you climb, pay in the controls that rung needs. Then climb only when your evals show the gain is worth it.
Where Anthropic and Cognition agree, and where they don’t
In June 2025, two well-known agent builders published opposing posts a day apart. Cognition’s Walden Yan wrote “Don’t Build Multi-Agents.” Anthropic described how it built its multi-agent research system. Read together, they agree more than their titles suggest.
Cognition argues from two principles: “Share context” and “Actions carry implicit decisions.” In plain terms, sub-agents working from partial context can make decisions that quietly conflict. Its conclusion: “The simplest way to follow the principles is to just use a single-threaded linear agent.”
Anthropic reports that its multi-agent setup beat a single agent by 90.2% on its internal research eval. The price is tokens. Agents use about 4 times the tokens of chat, and multi-agent systems about 15 times. Token usage alone explained 80% of performance variance on the BrowseComp benchmark. Anthropic also says multi-agent is a poor fit for “most coding tasks” and for work that needs shared context.
So both agree that tasks needing shared context break multi-agent setups. They differ on whether parallel breadth justifies roughly 15 times the tokens. For broad research, where sub-tasks are independent, Anthropic’s numbers make a case. For tightly coupled work like coding, both posts point toward a single agent. Keep in mind that the 90.2% comes from Anthropic’s own internal eval, not an external benchmark.
| Cognition | Anthropic | |
|---|---|---|
| Position | Don’t build multi-agents; use a single-threaded linear agent | Multi-agent works for broad research tasks |
| Main argument or evidence | Sub-agents with partial context make conflicting decisions | +90.2% on its internal research eval, at about 15x chat tokens |
| Where they agree | Tasks that need shared context break multi-agent setups | Poor fit for most coding tasks and shared-context work |
One talk worth your time is “12-Factor Agents: Patterns of reliable LLM applications.” Dex Horthy of HumanLayer gave it at AI Engineer World’s Fair 2025.
How to apply it: ten design controls
Reliability comes from controls you can point to in code. Air Canada shows what happens without them. In February 2024, a British Columbia tribunal ruled in Moffatt v. Air Canada (2024 BCCRT 149), a case about the airline’s chatbot (Dentons). The bot had told a customer he could claim a bereavement fare retroactively, contrary to the airline’s policy. The tribunal found negligent misrepresentation and awarded a total of CAD $812.02.
Air Canada had argued that the chatbot was “a separate legal entity that is responsible for its own actions.” The tribunal called that “a remarkable submission” and rejected it. The sum was small. The principle isn’t: you own whatever your agent says and does. These ten controls make that ownership manageable.
Constrain what the agent can do
- Narrow the scope. Give each agent one job and a clear definition of done. A bot that answers fare questions should answer from the fare policy, or hand off to a person.
- Wrap the model in a deterministic shell. Let code own routing, state, loop limits, and exit conditions. OpenAI’s guide recommends exit conditions such as a maximum number of turns, which targets MAST’s repetition and termination failures.
- Use structured outputs, then validate them. Structured-output modes, such as OpenAI’s Structured Outputs, make responses adhere to a supplied JSON Schema. Still validate the content and handle refusals, because schema-valid isn’t the same as correct.
- Grant least-privilege tools and authorize downstream. Give the agent only the tools and permissions the task needs. OWASP’s Excessive Agency guidance says to “implement authorization in downstream systems rather than relying on an LLM to decide if an action is allowed.”
- Make actions idempotent or reversible. An idempotent request “can be retransmitted or retried with no additional side effects,” per the AWS Builders’ Library. Prefer drafts, soft deletes, and staged changes, so a bad action costs an undo instead of a recovery.
- Separate environments. The agent that experiments should never hold production credentials. Automatic dev and prod database separation led the list of fixes Replit announced after its incident.
- Require human approval for irreversible actions. OpenAI’s guide calls for human intervention on “actions that are sensitive, irreversible, or have high stakes.” Make the approval a gate in code, not a sentence in the prompt.
Here’s what controls 2, 4, 5, and 7 look like in code. It’s an illustrative sketch, not any specific framework’s API.
# Illustrative sketch, not a specific framework's API.
MAX_TURNS = 12
IRREVERSIBLE = {"issue_refund", "delete_record", "send_email"}
def run_agent(task, llm, tools, approvals):
history = [task]
for _ in range(MAX_TURNS): # control 2: code owns the loop
step = llm.next_action(history) # schema-validated output
if step.kind == "finish":
return step.result
if step.tool not in tools: # control 4: granted tools only
raise PermissionError(step.tool)
if step.tool in IRREVERSIBLE: # control 7: a gate in code
if not approvals.request(step): # a person approves or rejects
history.append(f"{step.tool} rejected by reviewer")
continue
# control 5: an idempotency key makes a retried call safe
result = tools[step.tool](**step.args, idempotency_key=step.id)
history.append(result)
raise TimeoutError("Turn limit reached; escalate to a person")
Notice what the prompt doesn’t do here. It doesn’t enforce the turn limit, the tool list, or the approval. Code does.
Prove it works, and keep it working
- Gate releases on evals with pass^k. Run each eval task several times and track pass^k, the chance that every trial succeeds, alongside single-run accuracy. Build the set from real failures; we cover the setup in how to design evals for AI agents.
- Trace every step. Record each model call, tool call, input, and output so you can replay a failure. Anthropic runs full production tracing on its research system, and the OpenTelemetry GenAI semantic conventions (still in development) offer a vendor-neutral schema.
- Bound retries and plan fallbacks. Google SRE advises limiting retries per request, adding randomized exponential backoff, and degrading gracefully (Google SRE). Exercise the fallback path regularly, because “the code path you never use is the code path that (often) doesn’t work.”
Caveats: when complexity earns its keep
Complexity has to earn its place, and sometimes it does.
Breadth-heavy tasks are the clearest case. When a question splits into independent threads, parallel sub-agents can cover more ground than one agent working in sequence. Anthropic’s research results support that.
The eval delta is the deciding number. If a more complex design wins on your evals by enough to cover its extra cost and risk, take it. If you can’t measure the delta, you can’t justify the rung.
The compounding math is an illustration, not a law. It assumes independent steps, and real agent steps aren’t independent. Verification steps can catch errors, so a well-designed long chain can beat the curve. That supports the thesis rather than undercutting it: the checks are the design.
Some benchmark numbers are dated. Kapoor and τ-bench used 2024-era models, and METR’s Claude 3.7 Sonnet example comes from 2025. Absolute scores will move as models improve. The 2026 Rabanser results suggest the reliability gap hasn’t closed, though that’s one research group’s measurement.
Reader questions
Won’t better models just fix this?
Better models help, but they won’t close the gap on their own. The MAST authors say better base models alone “will be insufficient,” because the failures come from coordination and design. Rabanser et al. found that capability gains have produced only small reliability gains. And no model upgrade turns a prompt instruction into a permission check. Replit’s agent had the instruction. It didn’t have the control.
Isn’t multi-agent where everything is heading?
For some workloads, maybe. But its strongest advocates scope it carefully. Anthropic calls multi-agent a poor fit for most coding tasks and for work that needs shared context. OpenAI recommends getting the most out of a single agent first. Gall’s law suggests the same path: a complex system that works grows from a simple one that worked. Multi-agent is a rung you earn, not a starting point.
We already built the complex version. Now what?
Don’t rewrite everything at once. Instrument first: add tracing and measure pass^k on the tasks that matter most. Then try removing one hop at a time, replacing an agent hand-off with deterministic code and comparing on your evals. Sometimes Gall is right that the system “cannot be patched up,” and a simple rebuild beats incremental repair. If you’d rather not do it alone, here’s how we approach untangling an over-engineered AI system.
Where to start
This week, list every action your agent can take. Include every tool call, write, delete, message, and payment. Mark the ones you can’t undo. For each irreversible action, put an approval gate or an undo path in code, not in the prompt. If an action has neither, remove the tool until it does.
If you’d like help designing those controls, talk to Moonlight AI.
- ai-agents
- reliability
- agent-architecture
- ai-engineering