All posts

Graph Engineering for AI Agents: Design the Shape First

How to design multi-agent workflows as graphs: remove waits that move no data, give nodes schemas, verify against anchors, and know when one agent wins.

Misha Druzhinin
Glowing blue nodes joined by thin lines into a twisting mesh against a dark navy background.

A multi-agent system is usually sketched as a sequence of boxes and arrows. That sketch quietly becomes the execution plan. Every arrow turns into a wait, including arrows that carry nothing.

rvaniaaa’s X article “Graph Engineering with Claude” names the fix. Their claim is that the model was never the bottleneck; the shape of the work was. We’d say “rarely” rather than “never,” but the point stands. Design that shape as a dependency graph, and independent jobs stop queuing behind each other.

This post is our take on the idea. We checked the article’s claims against primary sources and corrected two of them. We rewrote its code against the real Claude Code workflow API. And we added what we’d tell a team first: verifiers need anchors, graphs fail quietly, and sometimes one agent is the better design.

Key Takeaways

  • A graph has nodes (one job each, with a fixed output schema) and edges (places where one node’s output feeds another). An arrow that moves no data is just a wait.
  • Anthropic’s research feature fans out to parallel subagents and synthesizes their results. It beat single-agent Claude Opus 4 by 90.2% on an internal eval, using about 15 times a chat’s tokens (Anthropic, 2025).
  • Same-model reviewers still favor their own model’s work. Verification needs an anchor: a test that ran, a record that exists.
  • Before building a graph, find two jobs that don’t depend on each other. If there aren’t any, one agent is the right design.

What graph engineering means

Graph engineering is deciding the dependency structure of agent work before writing any prompts. The output is a map of who waits on whom, and why.

The vocabulary is small:

  • A node is a job with one input and one typed output. In the support example below, “classify the ticket” and “look up the account” are separate nodes. Merge them, and neither can be tested or rerun alone.
  • An edge is a data dependency. Draw one only where a node consumes another node’s output.

None of this is new theory. In “Building effective agents” (December 2024), Anthropic separates workflows, which follow “predefined code paths,” from agents that direct their own process. It describes building blocks including parallel sectioning, voting, orchestrator-workers, and evaluator loops.

What’s new in 2026 is tooling. Claude Code’s dynamic workflows have Claude write a JavaScript orchestration script for a task, then run it in the background. Per the docs, the script “holds the loop, the branching, and the intermediate results itself.” Claude’s own context only receives the final answer.

The ideas don’t depend on that tool. They apply equally to LangGraph, a job queue, or plain asyncio. Our examples use Claude Code because its API is compact and the source article targets it.

Step 1: Remove arrows that carry no data

The source article’s first move is a test you can run on paper. For every arrow in your workflow, ask whether data actually crosses it. rvaniaaa calls the arrows that fail “fake edges.”

Take a support-reply agent built as a chain. It reads the ticket, classifies the intent, looks up the customer’s account, searches past incidents, then drafts a reply. Now check each arrow. The account lookup needs the customer ID from the ticket, not the intent label. The incident search needs the ticket text, not the account record. Only the draft needs everything.

A support-reply chain redrawn as a graph Top: a chain where the ticket is classified, then the account is looked up, then past incidents are searched, then a reply is drafted, each waiting for the previous step. Bottom: the same work as a graph. Classify, account lookup, and incident search each depend only on the ticket and run in parallel. Only the draft reply waits, because it needs all three results. Chain: each step waits for the last Ticket Classify Account Incidents Draft reply Graph: only real data dependencies Ticket Classify Account Incidents Draft reply
The same support-reply work as a chain and as a graph. Three lookups depend only on the ticket. Original diagram.

Redrawn, three lookups run at once, and the reply waits for the slowest of them rather than all three in turn. The test also works in reverse. If every arrow in a workflow carries data, there’s nothing to parallelize, and the chain is already the right design.

Step 2: Give every node a contract

Whatever reads a node’s output is usually code or another agent, not a person. A JSON field can be checked with an if. A paragraph needs another model call to interpret, and that call can be wrong.

In Claude Code workflows, passing a schema to agent() makes the subagent return JSON in that shape instead of prose (Claude Code docs):

const INTENT = {
  type: 'object',
  properties: {
    intent:     { type: 'string', enum: ['billing', 'bug', 'how-to', 'cancel'] },
    confidence: { type: 'number' },
    evidence:   { type: 'string' },  // the sentence in the ticket that decided it
  },
  required: ['intent', 'confidence', 'evidence'],
};

const label = await agent(`Classify this support ticket:\n${ticket}`, { schema: INTENT });

A schema also makes a node testable in isolation. It behaves like a function with a return type, so you can grade it the same way. Our guide to designing evals for AI agents covers how to build that test set from real failures.

Step 3: Use the diamond shape

When work is wide, one shape keeps recurring. Scope the work into items. Fan out one worker per item. Check each result. Merge in plain code. Then hand the survivors to a single synthesis step. The source article calls this the diamond.

Claude’s research feature uses the fan-out and synthesis halves of it. Anthropic describes an “orchestrator-worker pattern, where a lead agent coordinates the process while delegating to specialized subagents that operate in parallel” (Anthropic, 2025). An Opus 4 lead with Sonnet 4 subagents beat single-agent Opus 4 by 90.2% on Anthropic’s internal research eval.

The cost is real. In the same post, a single agent used about 4 times the tokens of a chat. The multi-agent system used about 15 times. Dividing those averages puts multi-agent at roughly four times a single agent. That’s our arithmetic, not a same-task measurement.

The diamond: scope, fan out, verify, merge, synthesize A scope step lists the work items. Four workers process items in parallel. Each result goes to a verify step on separate context, which tries to refute it against an anchor such as a test. A merge step in plain code counts and deduplicates the survivors without using a model. A final synthesize step writes one report. Scope list the work items Worker Worker Worker Worker Verify each result separate context, try to refute Merge in code count, dedupe, no model Synthesize one report from survivors Verify checks against an anchor: a test that ran, a record that exists
The diamond. Breadth comes from the fan-out; trust comes from the verify step and its anchor. Original diagram.

Here’s the diamond applied to a docs site. It finds every claim that cites a link, then checks that the linked page supports it. The cited page is the anchor. In Claude Code, you’d typically describe this in a sentence and let Claude generate the script. Treat that script as a design document. It’s the most precise description of what the run will do.

export const meta = {
  name: 'citation-audit',
  description: 'Check that every cited link in the docs supports its claim',
  phases: [{ title: 'Extract' }, { title: 'Check' }, { title: 'Report' }],
};

const CLAIMS = {
  type: 'object',
  properties: { claims: { type: 'array', items: {
    type: 'object',
    properties: { claim: { type: 'string' }, url: { type: 'string' } },
    required: ['claim', 'url'],
  } } },
  required: ['claims'],
};
const VERDICT = {
  type: 'object',
  properties: { supported: { type: 'boolean' }, why: { type: 'string' } },
  required: ['supported', 'why'],
};

const pages = args.pages; // scoped before the run, e.g. 20 pages to start

const results = await pipeline(
  pages,
  // Fan out: one extractor per page.
  (_, page) => agent(`List every factual claim in ${page} that cites a link.`,
    { phase: 'Extract', schema: CLAIMS }),
  // Verify: one checker per claim, given the claim and the link, nothing else.
  (found, page) => found === null ? null : parallel(found.claims.map(c => () =>
    agent(`Open ${c.url}. Try to show it does NOT support this claim: "${c.claim}"`,
      { phase: 'Check', schema: VERDICT })))
    .then(verdicts => ({
      page,
      unsupported: found.claims
        .map((c, i) => ({ ...c, why: verdicts[i]?.why }))
        .filter((_, i) => verdicts[i] && !verdicts[i].supported),
      unchecked: found.claims.filter((_, i) => verdicts[i] === null),
    })),
);

// Merge in plain code, and count what came back before trusting it.
const failed = results.filter(r => r === null).length;
if (failed) log(`WARNING: ${failed} of ${pages.length} pages returned no result`);

return agent(`Write one report of unsupported claims, grouped by page.
List unchecked claims separately, and say that ${failed} pages failed to load.
${JSON.stringify(results.filter(Boolean))}`, { phase: 'Report' });

Notice that checking happens per page, inside the pipeline, before the merge. Streaming lets a page’s claims be checked as soon as that page is extracted. When results overlap across items, flip that order. Deduplicate first, so you don’t pay to check the same claim twice.

Step 4: Verify on separate context, against an anchor

A fan-out multiplies output. It doesn’t make any of that output more trustworthy. That takes a check by something other than the producer.

Research supports that rule, with limits. Huang et al. found that “LLMs struggle to self-correct their responses without external feedback” (Huang et al., ICLR 2024). Their tests ran self-correction inside the same conversation.

Panickssery et al. go further, and the finding is less comfortable. On summarization tasks, models judging outputs in a separate prompt still scored their own generations higher than others’ (Panickssery et al., 2024). Human annotators rated the same outputs as equal. So a fresh context removes the shared reasoning, but not the model’s taste for its own work.

The Bun rewrite shows the practitioner version of this rule. Each implementer had “2 or more adversarial reviewers” (Sumner, 2026). The reviewer’s only job, per Sumner, was to “find bugs & reasons why the code does not work.”

In Claude Code, every agent() call already runs as its own subagent with its own context. So separation is the default, and the prompt is where you lose it. Paste the worker’s reasoning into the verifier, and the two calls share a mind again. Pass only the claim and the evidence.

Because same-model reviewers keep some bias, a panel of them can still agree on something wrong. The source article’s answer is anchors, “nodes that can not be argued with.” Examples include a test suite that ran, a row in the database, or money in the account. Anchors are the defense against Goodhart’s law: “when a measure becomes a target, it ceases to be a good measure” (Strathern, 1997).

Bun’s anchor was its test suite. Sumner merged only once 100% of it passed in CI on all platforms. He also “manually verified the tests were in fact running and not being skipped.” An anchor nobody has checked is just another opinion. Where your workflow has no anchor, it’s worth also testing a different model as the verifier.

Step 5: Stream by default, and use a barrier only when you need one

How stages connect sets the wall-clock time. Claude Code offers two primitives (Claude Code docs):

Primitive Next stage starts when Use it for
pipeline() Each item finishes its own previous stage The default for multi-stage work
parallel() Every task in the batch has finished Steps that need the whole set at once

Here’s a hypothetical. Run the support-reply graph from Step 1 over a backlog of 50 tickets. Most account lookups return in a second, but a few hit a slow CRM page and take 15. With a barrier after the lookup stage, no reply drafting starts until the slowest lookup returns. With a pipeline, each ticket’s draft starts as soon as its own lookups finish.

The source article lists the cases where a barrier earns its wait:

  • Deduplicating across all results before an expensive step.
  • Skipping a whole stage when the total is zero.
  • Comparing each result against the rest of the set.

Anything else, such as a map or filter between stages, can run inside a pipeline stage.

Step 6: For work of unknown size, loop until dry

Discovery tasks rarely announce their size. Hunting flaky tests, for instance: fixing one often unmasks the next. Library migrations behave the same way, as each converted module turns up new call sites. Here the graph needs a cycle, with explicit exit conditions.

The pattern is to keep running finders until two rounds in a row find nothing new. The Claude Code docs describe the same stopping rule. The source article flags the detail that makes it converge. Deduplicate new finds against everything already seen, including findings the verifier rejected. Track only confirmed results, and rejected findings come back every round.

The source pairs the dry-round rule with an iteration cap and a budget. Our version below keeps the round cap and adds one guard: a round where finders failed doesn’t count as dry. FINDERS, BUGS, BUG_VERDICT, key, and MAX_ROUNDS are defined elsewhere.

const seen = new Set();
const confirmed = [], unverified = [];
let dryRounds = 0, rounds = 0;

while (dryRounds < 2 && rounds < MAX_ROUNDS) {
  rounds++;
  const raw = await parallel(FINDERS.map(f => () => agent(f.prompt, { schema: BUGS })));
  const failed = raw.filter(r => r === null).length;
  if (failed) log(`Round ${rounds}: ${failed} of ${FINDERS.length} finders failed`);

  const fresh = [...new Map(raw.filter(Boolean).flatMap(r => r.bugs)
    .filter(b => !seen.has(key(b))).map(b => [key(b), b])).values()];  // dedupe within the round too
  if (!fresh.length) { if (!failed) dryRounds++; continue; }  // a failed round isn't a dry one
  dryRounds = 0;
  fresh.forEach(b => seen.add(key(b)));  // seen, not confirmed

  const verdicts = await parallel(fresh.map(b => () =>
    agent(`Is this a real bug? Try to refute it: ${b.desc}`, { schema: BUG_VERDICT })));
  confirmed.push(...fresh.filter((_, i) => verdicts[i]?.real));
  unverified.push(...fresh.filter((_, i) => verdicts[i] === null));
}
return { confirmed, unverified, rounds };

Where graphs break

Failure in a chain is loud: the run stops. Failure in a graph can be quiet, because the merge step still produces a report.

Multi-agent failures are well documented. Cemri et al. annotated over 1,600 traces from seven multi-agent frameworks (Cemri et al., 2025). They found 14 failure modes in three groups: system design, inter-agent misalignment, and task verification. Three further patterns, which the source article also covers, are specific to graph-shaped work.

The merge can’t read everything. Say a sweep over 800 files returns 800 summaries. The synthesis agent can’t hold them all in context. Merge in tiers: summarize groups of a few dozen, then synthesize from those group summaries.

Shared resources are edges too. Two enrichment agents that call the same CRM API will trip its rate limit together. Two refactoring agents can both edit a shared utilities file. Nothing in either prompt reveals the dependency. It lives in the resource.

The Bun rewrite hit this about two minutes in. One agent ran git stash, another ran git stash pop, and then came git reset HEAD --hard. Sumner’s summary: “They were stepping on each other!”

The source article says Bun’s fix was a worktree per agent. It wasn’t. Sumner wrote that per-agent worktrees would exhaust disk space because Bun’s repository is too big. Instead, he banned git stash and git reset. He later ran four workflows at once, each in its own worktree, with 16 agents sharing each worktree.

So list every shared resource, then isolate it or write rules for it. When the repository is small enough, Claude Code’s isolation: 'worktree' option gives each agent its own copy.

Failures disappear in the merge. In Claude Code, agent() resolves to null when it’s stopped or hits an unrecoverable API error. The common .filter(Boolean) then removes it silently. A report built from 197 of 200 results doesn’t say so unless you make it. Both code examples above count nulls and log them for that reason.

When a graph is the wrong design

Fan-out buys throughput, and by our arithmetic on Anthropic’s averages it roughly quadruples token spend. It doesn’t make any single worker’s judgment better. When the work is narrow, you pay that cost for no speedup.

Cognition’s “Don’t Build Multi-Agents” (June 2025) makes the counter-case. Its principles: “Share context, and share full agent traces, not just individual messages,” and “Actions carry implicit decisions, and conflicting decisions carry bad results.” Agents working in parallel on coupled pieces make conflicting assumptions. Anthropic, for its part, recommends “finding the simplest solution possible” before adding complexity.

The source article lists cases where a graph doesn’t pay. We’d put them this way:

  • One-file changes. Orchestration overhead outweighs any parallel speedup.
  • Work you want to approve step by step. A graph is built to run unattended. Checkpoints after every step defeat that.
  • Open-ended investigation. When you don’t yet know the question, you want one agent you can redirect.
  • Tightly coupled steps. When each step reads the last one or shares design decisions with it, parallel workers will conflict.

rvaniaaa’s fake-edge test from Step 1 settles it. If no two jobs are independent, keep the single agent. We made the broader argument in Complexity Is Easy. Reliability Takes Design.

What the Bun rewrite tells you about cost

The Bun port is one of the largest public examples of this style of work. Jarred Sumner reports translating 535,496 lines of Zig into Rust in 11 days, with about 50 Claude Code workflows (Sumner, 2026). At peak, four workflows ran in separate worktrees with 16 agents each, about 64 at once. Pre-merge usage was “around $165,000 at API pricing.” Sumner discloses that Anthropic acquired Bun in December 2025.

The source article puts the output at “over a million lines of Rust.” Bun’s own unsafe-code statistics imply about 780,000 lines. Anthropic’s announcement of dynamic workflows says 750,000.

The rewrite also drew criticism. Zig creator Andrew Kelley objected to shipping that much code without line-by-line human review (Kelley, 2026). He rejected the argument that a good test suite is enough to catch everything.

Both views deserve weight. The graph wrote the code. What made Sumner willing to merge it was the anchor: a test suite he’d confirmed was really running.

Few teams start where Bun did. The project already had a strong test suite, an engineer who ran the effort full time, and a budget for six-figure token bills.

Common questions about graph engineering

Is graph engineering just multi-agent systems with a new name?

Mostly, with a useful change of emphasis. “Multi-agent” counts the agents. “Graph” describes the dependencies between their jobs. The dependencies tell you where parallel work is safe and where checks belong.

Do I need Claude Code workflows to do this?

No. The patterns work in LangGraph, a task queue, or a few hundred lines of asyncio. Claude Code’s workflows are convenient because Claude writes the script and intermediate results stay out of your session. Choose what your team can debug.

How much more does a graph cost than a single agent?

Anthropic reports single agents at about 4 times a chat’s tokens and its multi-agent system at about 15 times. Dividing those gives roughly 4 times, though that’s our arithmetic across different tasks, not a measurement. Each extra worker adds its own token spend. The Claude Code docs suggest using a smaller model “for stages that don’t need the strongest one.” Keep the strongest model for the verify and synthesize steps, where a wrong call costs the most.

Where to start

Graph engineering is a design habit first and a tool second. Remove arrows that move no data. Give every node a schema. Verify against anchors, not opinions. Count what comes back. And keep the single agent when the work isn’t wide.

If you run multi-agent systems that are slow, expensive, or quietly wrong, Moonlight AI can audit and fix them.

Your step for this week: pick one agent workflow and mark every arrow that carries no data. Redraw it, then run both versions on the same small eval set. Compare quality, token cost, and wall-clock time. Keep the graph only if it wins on quality per dollar.