All posts

Complexity Is Easy. Reliability Takes Design.

Adding agents, tools, and autonomy doesn't make AI systems reliable. Here's what the evidence shows, and the ten design controls that make agents dependable.

Misha Druzhinin
Rows of brass gears, dials, and linkages packed inside a historic tide-predicting machine, seen from above.

In July 2025, an AI coding agent deleted a production database in the middle of a code freeze. SaaStr’s Jason Lemkin was about 12 days into an experiment with Replit’s agent. He had told it not to change anything without approval, and he had declared a code and action freeze. The agent deleted the database anyway, The Register reported. It held records for about 1,200 executives and 1,190 companies. The instruction was a request. It wasn’t a control.

Then the agent made things worse. It fabricated data and claimed a rollback was impossible, according to Fortune. The data was later recovered. Replit CEO Amjad Masad’s verdict: “Unacceptable and should never be possible.”

Look at what Replit announced next. Automatic separation of development and production databases. Staging environments, in the works. A planning-only chat mode. Better backup and rollback. None of those is a better prompt. Every one is a design control.

That’s the argument of this post. Adding agents, tools, and autonomy is easy, and it looks impressive on stage. Making the result dependable takes design, and the evidence suggests design does more for reliability than adding agents.

Key Takeaways

  • Each step in an agent chain multiplies risk. In an illustration with independent 95%-reliable steps, 20 steps succeed end to end only about 36% of the time.
  • In the MAST study, multi-agent frameworks failed 41% to 86.7% of the time, mostly from design and coordination problems.
  • Simpler often wins: on HumanEval, a simple baseline scored 93.2% for $2.45, while a tree-search agent scored 88.0% for $134.50.
  • Reliability comes from controls in code: narrow scope, least-privilege tools, reversible actions, human approval for irreversible steps, and repeated-trial evals.

The conventional view: more agents, more capability

The common assumption is that capability scales with architecture. Add a planner agent, a researcher, a coder, and a critic, and the system should handle harder work. Give it more tools and longer autonomy, and it should need less supervision.

It’s an easy story to believe, because it demos well. Each agent maps to a role on an org chart, so the design feels familiar. A five-agent demo that finishes a task once looks a lot like a five-agent product.

Complexity is cheap to add. Its cost arrives later, in production, as failures that don’t reproduce and traces nobody can follow. Even Anthropic, writing about its own multi-agent system, warns that “the gap between prototype and production is often wider than anticipated” (Anthropic). The rest of this post explains why that gap exists and how to design across it.

Why complexity doesn’t buy reliability

Complexity doesn’t buy reliability because every step you add is another place to fail, and failures compound. Reliability engineers have modeled this for decades. The NIST/SEMATECH e-Handbook describes the series model. If a system fails when any one component fails, its reliability is the product of the components’ reliabilities. The model assumes components fail independently.

Here’s what that looks like as an illustration, assuming independent steps that are each 95% reliable. Ten steps succeed end to end about 60% of the time. Twenty steps succeed about 36% of the time. Raise each step to 99%, and 20 steps reach about 82%, while 50 steps fall to about 61%. Getting 50 steps to roughly 95% takes 99.9% reliability per step.

End-to-end success at 95% per step Area chart of end-to-end success for a chain of independent steps that are each 95% reliable: 1 step 95.0%, 5 steps 77.4%, 10 steps 59.9%, 15 steps 46.3%, 20 steps 35.8%, 30 steps 21.5%, 40 steps 12.9%, 50 steps 7.7%. Derived from the series reliability model in the NIST/SEMATECH e-Handbook; real agent steps are not fully independent. End-to-end success at 95% per step Illustration: independent steps, each 95% reliable 95% 1 step 77.4% 5 steps 59.9% 10 steps 46.3% 15 steps 35.8% 20 steps 21.5% 30 steps 12.9% 40 steps 7.7% 50 steps Source: Derived from NIST/SEMATECH e-Handbook, series model
Source: Derived from NIST/SEMATECH e-Handbook, series model.

Treat that curve as a shape, not a forecast. LLM steps aren’t independent: errors correlate, and a later check can catch an earlier mistake. But the direction holds. Longer chains and more hand-offs cost reliability unless something verifies the work along the way.

The best direct evidence on multi-agent systems comes from the MAST study (Cemri et al., NeurIPS 2025). The authors annotated more than 1,600 traces from seven open-source frameworks, including MetaGPT, ChatDev, AG2, and Magentic-One. Failure rates ranged from 41% to 86.7%. They sorted failures into 14 modes across three categories: system design, inter-agent misalignment, and task verification. System design was the largest category.

Most common multi-agent failure modes Lollipop chart of the six most common failure modes in the MAST study of seven open-source multi-agent frameworks, as a share of annotated failures: Step repetition 15.7%, Reasoning-action mismatch 13.2%, Unaware of termination 12.4%, Disobey task spec 11.8%, Incorrect verification 9.1%, Incomplete verification 8.2%. Source: Cemri et al., Why Do Multi-Agent LLM Systems Fail?, arXiv 2503.13657 v3, 2025. Most common multi-agent failure modes Share of annotated failures, MAST taxonomy Step repetition 15.7 Reasoning-actionmismatch 13.2 Unaware oftermination 12.4 Disobey taskspec 11.8 Incorrectverification 9.1 Incompleteverification 8.2 Source: Cemri et al., Why Do Multi-Agent LLM Systems Fail? (2025)
Source: Cemri et al., Why Do Multi-Agent LLM Systems Fail?, 2025.

The authors’ conclusion is blunt. Failures “arise from the challenges in organizational design and agent coordination rather than the limitations of individual agents.” Better base models alone “will be insufficient.” Design fixes helped but didn’t close the gap. In the paper’s first version, ChatDev’s score rose from 25.0% to 34.4% with better prompts, then to 40.6% with a new topology. The authors still called that “insufficiently low for real-world deployment.”

Notice how many of the top failure modes are control-flow problems. Step repetition and not knowing when to stop are things ordinary code handles with a loop counter and an exit condition. That’s less a model weakness than a missing design decision.

What the evidence shows

Across several research groups and benchmarks, the pattern repeats. Simple setups match complex ones on accuracy, and consistency lags well behind capability.

Start with cost. In “AI Agents That Matter,” Kapoor et al. compared coding agents on HumanEval using GPT-4. Zero-shot prompting scored 89.6% for $1.93. A simple baseline the authors call warming scored 93.2% for $2.45. LATS, a tree-search agent, scored 88.0% for $134.50. Their verdict: “SOTA agents are needlessly complex and costly.” And: “for substantially similar accuracy, the cost can differ by almost two orders of magnitude.”

Cost to run HumanEval with GPT-4 (USD) Horizontal bar chart of the total cost in US dollars to run HumanEval with GPT-4 for six approaches, with accuracy: zero-shot 89.6% at $1.93, Warming 93.2% at $2.45, Retry 92.0% at $2.51, Reflexion 87.8% at $3.90, LDB 93.3% at $6.36, LATS 88.0% at $134.50. Source: Kapoor et al., AI Agents That Matter, 2024, Table A1. Cost to run HumanEval with GPT-4 (USD) Approach and accuracy on the left; total cost on the bars Zero-shot, 89.6% $1.93 Warming, 93.2% $2.45 Retry, 92.0% $2.51 Reflexion, 87.8% $3.9 LDB, 93.3% $6.36 LATS, 88.0% $134.5 Source: Kapoor et al., AI Agents That Matter (2024)
Source: Kapoor et al., AI Agents That Matter, 2024.

Then consistency. τ-bench (Yao et al., 2024) introduced pass^k: the chance that all k independent trials of the same task succeed, averaged across tasks. A gpt-4o function-calling agent scored a pass^1 of 61.2% on retail tasks and 35.2% on airline tasks. On retail, its pass^8 fell below 25%. So an agent that succeeds on most single tries succeeds on all eight tries less than a quarter of the time. In production, every user is another trial.

METR measured the same gap in task length. Kwa et al. found that a model’s 80% success time horizon is 4 to 6 times shorter than its 50% horizon. In the paper’s first version, Claude 3.7 Sonnet’s 50% horizon was about 59 minutes, and its 80% horizon about 15 minutes. Those are 2025-era models, but the ratio is the lesson. What a model can do sometimes is much longer than what it does reliably.

The newest work points the same way. Rabanser et al. (ICML 2026) scored 15 models on 12 reliability metrics spanning consistency, robustness, predictability, and safety. Their finding: “recent capability gains have only yielded small improvements in reliability.” One caveat: this paper shares authors with the Kapoor study. Treat the two as one research thread, not independent confirmation.

Analysts expect the consequences to show up in budgets. Gartner forecasts that over 40% of agentic AI projects will be canceled by the end of 2027. It cites “escalating costs, unclear business value or inadequate risk controls” (Gartner). That’s a forecast, not a measurement.

Reliability is a design property

Reliability doesn’t emerge from capability. You design it in, and you start simple. John Gall put it plainly in Systemantics (1975):

“A complex system that works is invariably found to have evolved from a simple system that worked. A complex system designed from scratch never works and cannot be patched up to make it work. You have to start over with a working simple system.”

Google’s SRE book makes the same point in operational terms: “Software simplicity is a prerequisite to reliability” (Google SRE, Simplicity). It separates essential complexity, which comes with the problem, from accidental complexity, which the solution adds. A multi-agent hand-off for a task one agent could handle is usually the accidental kind.

SRE also gives a frame for deciding how much reliability to buy. “100% is probably never the right reliability target,” and each extra increment may cost 100 times the previous one (Google SRE, Embracing Risk). The tool is an error budget: the amount of failure you agree to tolerate, which you then spend deliberately. Every agent, tool, and autonomous step draws from that budget.

The major vendors give the same advice. Anthropic’s “Building effective agents” defines workflows as “predefined code paths,” while agents “dynamically direct their own processes and tool usage.” It recommends finding “the simplest solution possible, and only increasing complexity when needed.” OpenAI’s practical guide to building agents says to get the most out of a single agent before splitting into many. Its advice: “Start small, validate with real users, and grow capabilities over time.”

Barry Zhang of Anthropic expands on when an agent is worth building in “How We Build Effective Agents,” from AI Engineer Summit 2025.

We think of that advice as a complexity budget. Each rung of the ladder below adds capability and risk. Before you climb, pay in the controls that rung needs. Then climb only when your evals show the gain is worth it.

The complexity budget A ladder of four rungs, from simplest at the bottom to most complex at the top. Each rung keeps every control from the rungs below it. Rung one, one LLM call, needs structured output with validation and a small eval set built from real failures. Rung two, a deterministic workflow, adds tracing on every step, bounded retries with fallbacks, and pass^k release gates. Rung three, a single agent with tools, adds least-privilege tools, reversible actions, sandboxed and separate environments, and human approval on irreversible actions. Rung four, a multi-agent system, adds shared context, one owner per decision, and an explicit token budget. An arrow up the side says to climb a rung only when evals show the gain is worth the added risk and cost. 4. Multi-agent system shared context, one owner per decision, an explicit token budget 3. Single agent with tools least-privilege tools, reversible actions, sandboxed and separate environments, human approval on irreversible actions 2. Deterministic workflow tracing on every step, bounded retries and fallbacks, pass^k release gates 1. One LLM call structured output and validation, a small eval set from real failures Climb a rung only when evals show the gain is worth the added risk and cost
The complexity budget: the controls each rung adds before you climb to the next. Each rung keeps every control below it. "One owner per decision" means a single agent makes each decision that others depend on (see Cognition below). Original framework diagram.

Where Anthropic and Cognition agree, and where they don’t

In June 2025, two well-known agent builders published opposing posts a day apart. Cognition’s Walden Yan wrote “Don’t Build Multi-Agents.” Anthropic described how it built its multi-agent research system. Read together, they agree more than their titles suggest.

Cognition argues from two principles: “Share context” and “Actions carry implicit decisions.” In plain terms, sub-agents working from partial context can make decisions that quietly conflict. Its conclusion: “The simplest way to follow the principles is to just use a single-threaded linear agent.”

Anthropic reports that its multi-agent setup beat a single agent by 90.2% on its internal research eval. The price is tokens. Agents use about 4 times the tokens of chat, and multi-agent systems about 15 times. Token usage alone explained 80% of performance variance on the BrowseComp benchmark. Anthropic also says multi-agent is a poor fit for “most coding tasks” and for work that needs shared context.

So both agree that tasks needing shared context break multi-agent setups. They differ on whether parallel breadth justifies roughly 15 times the tokens. For broad research, where sub-tasks are independent, Anthropic’s numbers make a case. For tightly coupled work like coding, both posts point toward a single agent. Keep in mind that the 90.2% comes from Anthropic’s own internal eval, not an external benchmark.

Cognition Anthropic
Position Don’t build multi-agents; use a single-threaded linear agent Multi-agent works for broad research tasks
Main argument or evidence Sub-agents with partial context make conflicting decisions +90.2% on its internal research eval, at about 15x chat tokens
Where they agree Tasks that need shared context break multi-agent setups Poor fit for most coding tasks and shared-context work

One talk worth your time is “12-Factor Agents: Patterns of reliable LLM applications.” Dex Horthy of HumanLayer gave it at AI Engineer World’s Fair 2025.

How to apply it: ten design controls

Reliability comes from controls you can point to in code. Air Canada shows what happens without them. In February 2024, a British Columbia tribunal ruled in Moffatt v. Air Canada (2024 BCCRT 149), a case about the airline’s chatbot (Dentons). The bot had told a customer he could claim a bereavement fare retroactively, contrary to the airline’s policy. The tribunal found negligent misrepresentation and awarded a total of CAD $812.02.

Air Canada had argued that the chatbot was “a separate legal entity that is responsible for its own actions.” The tribunal called that “a remarkable submission” and rejected it. The sum was small. The principle isn’t: you own whatever your agent says and does. These ten controls make that ownership manageable.

Constrain what the agent can do

  1. Narrow the scope. Give each agent one job and a clear definition of done. A bot that answers fare questions should answer from the fare policy, or hand off to a person.
  2. Wrap the model in a deterministic shell. Let code own routing, state, loop limits, and exit conditions. OpenAI’s guide recommends exit conditions such as a maximum number of turns, which targets MAST’s repetition and termination failures.
  3. Use structured outputs, then validate them. Structured-output modes, such as OpenAI’s Structured Outputs, make responses adhere to a supplied JSON Schema. Still validate the content and handle refusals, because schema-valid isn’t the same as correct.
  4. Grant least-privilege tools and authorize downstream. Give the agent only the tools and permissions the task needs. OWASP’s Excessive Agency guidance says to “implement authorization in downstream systems rather than relying on an LLM to decide if an action is allowed.”
  5. Make actions idempotent or reversible. An idempotent request “can be retransmitted or retried with no additional side effects,” per the AWS Builders’ Library. Prefer drafts, soft deletes, and staged changes, so a bad action costs an undo instead of a recovery.
  6. Separate environments. The agent that experiments should never hold production credentials. Automatic dev and prod database separation led the list of fixes Replit announced after its incident.
  7. Require human approval for irreversible actions. OpenAI’s guide calls for human intervention on “actions that are sensitive, irreversible, or have high stakes.” Make the approval a gate in code, not a sentence in the prompt.

Here’s what controls 2, 4, 5, and 7 look like in code. It’s an illustrative sketch, not any specific framework’s API.

# Illustrative sketch, not a specific framework's API.
MAX_TURNS = 12
IRREVERSIBLE = {"issue_refund", "delete_record", "send_email"}

def run_agent(task, llm, tools, approvals):
    history = [task]
    for _ in range(MAX_TURNS):               # control 2: code owns the loop
        step = llm.next_action(history)      # schema-validated output
        if step.kind == "finish":
            return step.result
        if step.tool not in tools:           # control 4: granted tools only
            raise PermissionError(step.tool)
        if step.tool in IRREVERSIBLE:        # control 7: a gate in code
            if not approvals.request(step):  # a person approves or rejects
                history.append(f"{step.tool} rejected by reviewer")
                continue
        # control 5: an idempotency key makes a retried call safe
        result = tools[step.tool](**step.args, idempotency_key=step.id)
        history.append(result)
    raise TimeoutError("Turn limit reached; escalate to a person")

Notice what the prompt doesn’t do here. It doesn’t enforce the turn limit, the tool list, or the approval. Code does.

Prove it works, and keep it working

  1. Gate releases on evals with pass^k. Run each eval task several times and track pass^k, the chance that every trial succeeds, alongside single-run accuracy. Build the set from real failures; we cover the setup in how to design evals for AI agents.
  2. Trace every step. Record each model call, tool call, input, and output so you can replay a failure. Anthropic runs full production tracing on its research system, and the OpenTelemetry GenAI semantic conventions (still in development) offer a vendor-neutral schema.
  3. Bound retries and plan fallbacks. Google SRE advises limiting retries per request, adding randomized exponential backoff, and degrading gracefully (Google SRE). Exercise the fallback path regularly, because “the code path you never use is the code path that (often) doesn’t work.”

Caveats: when complexity earns its keep

Complexity has to earn its place, and sometimes it does.

Breadth-heavy tasks are the clearest case. When a question splits into independent threads, parallel sub-agents can cover more ground than one agent working in sequence. Anthropic’s research results support that.

The eval delta is the deciding number. If a more complex design wins on your evals by enough to cover its extra cost and risk, take it. If you can’t measure the delta, you can’t justify the rung.

The compounding math is an illustration, not a law. It assumes independent steps, and real agent steps aren’t independent. Verification steps can catch errors, so a well-designed long chain can beat the curve. That supports the thesis rather than undercutting it: the checks are the design.

Some benchmark numbers are dated. Kapoor and τ-bench used 2024-era models, and METR’s Claude 3.7 Sonnet example comes from 2025. Absolute scores will move as models improve. The 2026 Rabanser results suggest the reliability gap hasn’t closed, though that’s one research group’s measurement.

Reader questions

Won’t better models just fix this?

Better models help, but they won’t close the gap on their own. The MAST authors say better base models alone “will be insufficient,” because the failures come from coordination and design. Rabanser et al. found that capability gains have produced only small reliability gains. And no model upgrade turns a prompt instruction into a permission check. Replit’s agent had the instruction. It didn’t have the control.

Isn’t multi-agent where everything is heading?

For some workloads, maybe. But its strongest advocates scope it carefully. Anthropic calls multi-agent a poor fit for most coding tasks and for work that needs shared context. OpenAI recommends getting the most out of a single agent first. Gall’s law suggests the same path: a complex system that works grows from a simple one that worked. Multi-agent is a rung you earn, not a starting point.

We already built the complex version. Now what?

Don’t rewrite everything at once. Instrument first: add tracing and measure pass^k on the tasks that matter most. Then try removing one hop at a time, replacing an agent hand-off with deterministic code and comparing on your evals. Sometimes Gall is right that the system “cannot be patched up,” and a simple rebuild beats incremental repair. If you’d rather not do it alone, here’s how we approach untangling an over-engineered AI system.

Where to start

This week, list every action your agent can take. Include every tool call, write, delete, message, and payment. Mark the ones you can’t undo. For each irreversible action, put an approval gate or an undo path in code, not in the prompt. If an action has neither, remove the tool until it does.

If you’d like help designing those controls, talk to Moonlight AI.