All posts

How to Design Evals for AI Agents: An Eight-Step Method

An eight-step method for designing AI agent evals: build tasks from real failures, pick graders, report pass^k, calibrate LLM judges, and gate releases in CI.

Misha Druzhinin
Blue and white fiber-optic strands glowing against a black background.

Plenty of agents that pass a demo fall apart in production. The model is rarely the whole problem. More often, nobody wrote down what “working” means and then checked it, repeatedly, on real tasks.

The survey data points the same way. LangChain’s “State of Agent Engineering” survey ran from November 18 to December 2, 2025, with 1,340 respondents (LangChain, 2025). Of those, 57.3% said they have agents in production. 89% have observability in place, but only 52.4% run offline evals. The audience is self-selected from LangChain’s community, and 63% work in tech, so read it as a signal rather than a census. The signal is still clear: teams watch their agents far more than they test them.

This guide gives you an eight-step method, from the first task to a CI gate. It’s written for staff engineers and technical leaders who already ship software and now own an agent. Every measured number links to its source, and derived or hypothetical numbers are labeled. Much of the terminology comes from Anthropic’s “Demystifying evals for AI agents,” the most complete public write-up we know of. We cross-check it against independent research.

Key Takeaways

  • Build your first 20-50 eval tasks from real failures found in traces, and pair every “should” case with a “should not” case.
  • Use the cheapest grader that can decide each check: code first, a calibrated LLM judge for fuzzy criteria, humans for ground truth.
  • Report pass^k, the chance that all k trials succeed. At 75% per-trial success, three straight passes happen only about 42% of the time (Anthropic, 2026).
  • Treat your harness and graders as code under test. A broken grader can hide real capability or reward wrong answers.

Before you start

You need four things before you write a single eval. None of them is a platform.

  • Traces. Transcripts of agent runs from a prototype, a pilot, or production. Real inputs beat invented ones.
  • One domain expert who owns “good.” Someone who can read an agent run and call it pass or fail without convening a committee.
  • A sandboxed environment. Test accounts, seeded databases, and staging or mocked tools, so trials can’t touch real customers or each other.
  • A few days. The first useful suite is small, but reading traces takes time. Husain and Shankar report that on their projects, 60-80% of development time went to error analysis and evaluation (Husain & Shankar, 2026). That’s their practitioner experience, not a law, but it sets honest expectations.

Step 1: Define success as outcomes, constraints, and “should not” cases

An agent eval is a task, an isolated environment to run it in, and a grader that decides whether the run succeeded. In effect, it’s an executable spec. Before you test anything, write down what the agent must achieve, which rules it must respect, and what it must never do.

Split each task’s definition into three parts:

  1. Outcome. The end state that counts as done: a refund issued, a ticket routed, a file patched.
  2. Constraints. Limits on how the agent gets there, such as budget, latency, permitted tools, and business policy.
  3. “Should not” cases. Actions that fail the task regardless of outcome, like refunding above a limit or emailing the wrong customer.

The third part is the one teams skip. Anthropic’s guidance is blunt: “One-sided evals create one-sided optimization” (Anthropic, 2026). Test only that the agent issues refunds, and you’ll tune it to issue refunds eagerly. Pair every “should” task with a similar “should not” task.

Write criteria specific enough that two people would grade the same run the same way. “Gives a helpful response” fails that test. “Escalates any refund above the auto-approval limit to a human” passes it.

Step 2: Build the first task set from real failures

Your first tasks should come from failures you’ve actually seen, not from a whiteboard session. Real failures are specific, they’re the ones users hit, and they come with inputs you can replay.

Start with error analysis. Husain and Shankar recommend reviewing at least 100 diverse traces, starting with 30 yourself (Husain & Shankar, 2026). Write a short note on what went wrong in each one, a step they call open coding. Then group those notes into a handful of recurring failure modes (axial coding). The counts per mode tell you where to spend effort first.

Then turn failure modes into tasks. Anthropic suggests that “20-50 simple tasks drawn from real failures is a great start” (Anthropic, 2026). That’s fewer than most teams expect. A small suite you actually read beats a large one nobody opens.

Mix sources as the suite grows. LangSmith documents converting production traces into dataset examples (LangSmith docs). OpenAI recommends adding expert-written cases to production data (OpenAI), which covers rare, expensive paths traffic hasn’t hit yet.

The agent eval lifecycle A loop of five stages: production traces feed error analysis, where failure modes are labeled. Failure modes become capability evals with a low pass rate. Once they pass reliably they graduate into the regression suite, which should pass close to 100 percent. The regression suite runs as a CI gate on every change. Shipped changes produce new production traces and the loop repeats. Production traces real agent runs Error analysis label failure modes Capability evals low pass rate at first Regression suite ~100% pass rate CI gate runs on every change Ship, observe in production, repeat
The agent eval lifecycle: failures found in production become capability evals, then regression tests. Original diagram.

Step 3: Pick the cheapest grader that can decide each check

Every check in a task needs a grader, and the right one is the cheapest grader that can decide it reliably. Anthropic describes three grader types with distinct trade-offs (Anthropic, 2026):

Grader Best for Strengths Weaknesses Cost and speed
Code checks Final state, tool arguments, formats, hard policy rules Fast, objective, reproducible Brittle; can reject valid answers that differ from the expected form Cheapest and fastest
LLM judge Tone, relevance, policy nuance, open-ended outputs Flexible; handles fuzzy criteria at scale Non-deterministic; needs calibration against humans Costs money per call; slower than code
Human review Ground truth, judge calibration, high-stakes samples The gold standard for quality Slow, expensive, hard to scale Most expensive and slowest

Default to code. Many agent tasks have a checkable end state: a database row, an API call, a file diff. Reach for a judge only when no deterministic check can settle the question. Save humans for calibrating judges and reviewing samples.

Most teams already blend these. In LangChain’s survey, 59.8% of respondents use human review and 53.3% use LLM-as-judge (LangChain, 2025). So the real question isn’t which grader your team uses. It’s which grader each individual check deserves.

Step 4: Grade the outcome first, the trajectory only where the path matters

Grade what the agent produced before you grade how it got there. Anthropic puts it plainly: “it’s often better to grade what the agent produced, not the path it took” (Anthropic, 2026). Agents find valid paths you didn’t anticipate, and strict path-matching punishes them for it.

τ-bench (tau-bench) is the reference example. It grades customer-service agents by comparing the final database state to an annotated goal state (Yao et al., 2024). It doesn’t care which order the agent looked things up in. It cares whether the right records changed, and only those.

The path does matter in three situations:

  • Cost and latency. An agent that reaches the right answer through dozens of redundant tool calls still has a problem.
  • Safety. Some intermediate actions are unacceptable even if the end state looks fine, such as reading data the user shouldn’t see.
  • Irreversible tools. A correct final answer doesn’t undo a payment, a deletion, or an email that already went out.

For those cases, add trajectory checks on top of outcome checks. Google’s Agent Development Kit scores tool trajectories with three match modes: EXACT, IN_ORDER, and ANY_ORDER (Google ADK docs). The default is EXACT with a threshold of 1.0, which is stricter than most tasks need. LangSmith separates final-response, single-step, and trajectory evaluations (LangSmith docs). Pick the loosest match that still catches the failure you care about.

Anatomy of an agent eval A task with inputs and success criteria is run through the agent harness, which contains the agent, its tools, and an isolated sandbox. Each run produces a transcript of tool calls and steps, and an outcome, the final state. Graders (code checks, an LLM judge, or human review) score the transcript or outcome as pass or fail per trial. Each task is repeated k times and reported as pass^k. Task inputs + success criteria Agent harness agent + tools, isolated sandbox Transcript tool calls, steps Outcome final state Graders code checks, LLM judge, human review Score pass or fail per trial Repeat each task k times, then report pass^k
Anatomy of an agent eval. Original diagram, adapted from the task, trial, and grader terminology in Anthropic's "Demystifying evals for AI agents".

For a practitioner’s take on the same discipline, Philipp Schmid of Google DeepMind gave this talk at the AI Engineer World’s Fair 2026.

Step 5: Run multiple trials and report pass^k, not pass@k

Agents are non-deterministic, so one run per task tells you very little. Run each task several times and report pass^k. τ-bench defines it as the chance that all k trials of a task succeed, averaged across tasks (Yao et al., 2024).

Its better-known cousin, pass@k, is the chance that at least one of k tries succeeds. That fits a human picking the best of several drafts. It flatters an agent that serves each customer once.

The gap between the two is large. Anthropic’s worked example: at 75% per-trial success, the chance of passing three trials in a row is about 42% (Anthropic, 2026). Extending that arithmetic, assuming independent trials, gives the numbers below. They’re our derivation for illustration, not measurements. pass^k falls to 56.3% at k=2, 23.7% at k=5, and 5.6% at k=10. pass@k moves the other way, reaching about 99.9999% at k=10.

pass^k at 75% per-trial success Line chart of pass^k for an agent with 75% per-trial success: 75% at k=1, 56.3% at k=2, 42.2% at k=3, 31.6% at k=4, 23.7% at k=5, 17.8% at k=6, 13.3% at k=7, 10% at k=8, 7.5% at k=9, 5.6% at k=10. Derived from 0.75 to the power k, following Anthropic's worked example. pass^k at 75% per-trial success Chance that all k trials succeed (illustrative math) 75% k=1 56.3% k=2 42.2% k=3 31.6% k=4 23.7% k=5 17.8% k=6 13.3% k=7 10% k=8 7.5% k=9 5.6% k=10 Source: Derived from Anthropic, Demystifying evals for AI agents (Jan 2026)
Source: Derived from Anthropic, Demystifying evals for AI agents, Jan 2026.

τ-bench shows the same curve with real agents. The best gpt-4o function-calling agent scored 61.2% pass^1 on the retail domain and 35.2% on airline. On retail, its pass^8 dropped below 25% (Yao et al., 2024). Those are 2024-era models, so the exact numbers have aged. The shape of the curve hasn’t.

This is the demo-versus-product gap, expressed as one number. A demo is effectively pass@k on a task you picked: someone quietly retries until it works. A product is pass^k across every customer, every day. When a leader asks “does it work?”, pass^k at a realistic k is the honest answer.

In practice, we’d start with five trials per task and raise k for tasks with real money or data behind them. Then look hardest at tasks that pass sometimes. We’d treat flaky tasks as the likeliest source of production incidents.

Step 6: Calibrate the LLM judge against human labels

An LLM judge is a model making predictions, so measure it like one, against your domain expert’s labels. Anthropic says judges “should be closely calibrated with human experts” (Anthropic, 2026). OpenAI advises scaling a judge only once it consistently agrees with human annotations (OpenAI).

The encouraging news is that judges can match human agreement. In Zheng et al.’s MT-Bench study, GPT-4 agreed with human experts 85% of the time, versus 81% between humans (Zheng et al., 2023). That was a pairwise setup with ties excluded, on 2023 chat outputs rather than agent runs.

The same study also measured how judges go wrong. When the two answers were swapped, GPT-4 kept its verdict only 65.0% of the time, GPT-3.5 46.2%, and Claude-v1 23.8%. Against a verbosity attack that padded answers with repetitive lists, GPT-3.5 and Claude-v1 failed 91.3% of the time. GPT-4 failed 8.7% of the time. Later work found that models also tend to prefer their own outputs (Wataoka et al., 2024). If you judge pairwise, swap the order and check that the verdict holds.

LLM judge position consistency Lollipop chart of position consistency for LLM judges on MT-Bench when the two answers are swapped: GPT-4 65.0%, GPT-3.5 46.2%, Claude-v1 23.8%. Source: Zheng et al., Judging LLM-as-a-Judge, NeurIPS 2023, Table 2. LLM judge position consistency Same verdict after swapping answer order (higher is better) GPT-4 65% GPT-3.5 46.2% Claude-v1 23.8% Source: Zheng et al., Judging LLM-as-a-Judge (NeurIPS 2023)
Source: Zheng et al., Judging LLM-as-a-Judge (NeurIPS 2023).

A calibration process that holds up looks like this:

  1. Use binary pass/fail verdicts. Husain and Shankar argue for binary judgments over Likert scales (Husain & Shankar, 2026). OpenAI also prefers pairwise or pass/fail judges (OpenAI). A “3 out of 5” hides the decision you actually need to make.
  2. Write one judge per failure mode. “Did the agent promise a refund it can’t issue?” is far easier to calibrate than “rate overall quality.”
  3. Label enough examples, and split them. Husain and Shankar suggest 100-200 labeled examples per failure mode, divided into train, dev, and test sets (Husain & Shankar, 2026). Few-shot examples come from train. You tune the prompt on dev and report agreement on test.
  4. Expect your criteria to move. Shankar et al. found that people refine their grading criteria while grading outputs, which they call criteria drift (Shankar et al., 2024). Re-label a sample whenever your definition of “good” shifts.

Publish the judge’s test-set agreement next to every score it produces. A judge score without its agreement number is an unverified claim.

Step 7: Treat the harness and the grader as code under test

Your eval infrastructure has bugs, and those bugs move your scores. Treat the harness and every grader as production code: review it, test it, and read its failures.

Anthropic’s CORE-Bench experience shows how big the effect can be. Its Opus 4.5 model initially scored 42%. After Anthropic fixed grading bugs and loosened the scaffold, meaning the prompts and tooling around the model, the score jumped to 95% (Anthropic, 2026). That’s Anthropic reporting on its own model, so weigh it accordingly. Still, the model stayed the same. The harness and the graders changed.

Public benchmarks share the problem. OpenAI audited 138 of SWE-bench Verified’s 500 tasks (27.6%), concentrating on tasks models often failed. It found that 59.4% of audited tasks had flawed tests that rejected functionally correct submissions (Epoch AI). Citing those flaws and training-data contamination, OpenAI stopped reporting the benchmark in early 2026.

OpenAI's audit of SWE-bench Verified Horizontal bar chart: SWE-bench Verified has 500 tasks; OpenAI audited 138 (27.6%), concentrating on tasks models often failed, and found 59.4% of those, about 82 tasks, had flawed tests that reject correct solutions. The 82 is derived from 59.4% of 138. Source: Epoch AI review of OpenAI's 2026 findings. OpenAI's audit of SWE-bench Verified Tasks in the benchmark, tasks audited, and audited tasks with flawedtests Tasks inbenchmark 500 Tasks audited 138 Flawed tests(~59.4%) 82 Source: Epoch AI, SWE-bench Verified review (2026)
Source: Epoch AI, SWE-bench Verified review, 2026. The 82 is derived: 59.4% of 138.

The lesson for your own suite: a failing eval is a claim, not a fact. When a task fails, read the transcript before you touch the prompt. Sometimes the agent is wrong, and sometimes the grader is. Unit-test each grader against known-good and known-bad transcripts, exactly as you’d test any other function. Auditing the harness and graders is also the first step when cleaning up an existing AI system.

Three hygiene rules close the most common gaps:

  • Isolate trials. Anthropic recommends running each trial in a clean environment (Anthropic, 2026). Shared state lets one trial’s side effects pass or fail the next one.
  • Give partial credit on multi-step tasks. Anthropic also suggests this, so a near-miss doesn’t look like a total failure.
  • Track cost and latency per task. Kapoor et al. argue for cost-controlled agent evaluation, plus holdout sets and reproducible setups (Kapoor et al., 2024). Latency belongs in the report too. In LangChain’s survey, it was the second-biggest barrier at 20%, behind quality at 32% (LangChain, 2025).

Step 8: Wire evals into the development loop

Evals pay off when they run on every change, not when someone remembers to run them. OpenAI’s guidance calls this eval-driven development: “Evaluate early and often” (OpenAI).

Run two kinds of suites. Capability evals measure what the agent can’t yet do reliably, so they start with low pass rates. Regression evals protect what already works, and Anthropic says they “should have a nearly 100% pass rate” (Anthropic, 2026). When a capability task passes consistently, it graduates into the regression suite.

Gate every change on the regression suite: prompt edits, model upgrades, tool changes, retrieval tweaks. Agents are noisy, so decide in advance how a failure blocks a merge. For example, gate on each task’s pass^k across several trials, or re-run a failed task once before blocking. Run capability evals on a schedule or before a release, since they’re expected to fail often.

Then close the loop with production. Online evals run the same graders on sampled live traffic. In LangChain’s survey, 37.3% of all respondents ran online evals, rising to 44.8% among teams with agents in production (LangChain, 2025). Meanwhile, 94% of those production teams had observability. Most teams can see what their agent does. Fewer score it.

Teams observe agents more than they evaluate them Grouped bar chart. All respondents vs teams with agents in production: observability 89% vs 94%, tracing 62% (detailed) vs 71.5% (full), online evals 37.3% vs 44.8%, not evaluating 29.5% vs 22.8%. Source: LangChain, State of Agent Engineering, survey Nov-Dec 2025, n=1,340. Teams observe agents more than they evaluate them All respondents (n=1,340) vs the production subgroup, % All respondents Agents in production 89 94 Observability 62 71.5 Tracing(detailed/full) 37.3 44.8 Online evals 29.5 22.8 Notevaluating Source: LangChain, State of Agent Engineering (survey Nov-Dec 2025)
Source: LangChain, State of Agent Engineering (survey Nov-Dec 2025).

New failures from production feed back into Step 2, and the loop repeats. Anthropic’s broader advice on agents applies here: add complexity only when it demonstrably improves outcomes (Anthropic, 2024). Your eval suite is how you find out whether it did.

Sayash Kapoor, lead author of the “AI Agents That Matter” paper cited above, covers agent evaluation in this AI Engineer Summit 2025 talk.

Worked example: one task, one grader, five trials

Here’s the method applied to a single customer-support task. It’s a hypothetical, simplified example, not any specific framework’s API. The agent can look up orders, issue refunds, and escalate to a human. Policy says refunds above $500 need human approval.

The task targets a “should not” case: a customer asks for more money back than the agent may refund on its own.

# Hypothetical, simplified task definition (not a specific framework's API)
id: refund-over-limit-001
input: "Order #4821 arrived broken. I paid $640 and want a full refund."
environment:
  seed_db: fixtures/order_4821.json  # order total: 640.00, status: delivered
policy:
  auto_refund_limit: 500.00
expected_outcome:
  refund_issued: false        # over the limit, so no direct refund
  escalation_created: true    # a human must approve it
should_not:
  - refund_above_limit_without_escalation
trials: 5

The grader checks the end state, τ-bench style. It fails the “should not” case outright, then compares the outcome to the goal.

def grade(final_state: dict, task: dict) -> bool:
    limit = task["policy"]["auto_refund_limit"]
    goal = task["expected_outcome"]
    refunds = final_state["refunds"]
    escalated = len(final_state["escalations"]) > 0

    # "Should not": any refund above the limit without escalation fails outright
    if any(r["amount"] > limit for r in refunds) and not escalated:
        return False

    # Outcome: compare the end state to the annotated goal
    return (bool(refunds) == goal["refund_issued"]
            and escalated == goal["escalation_created"])

Run each task five times, then estimate pass^k. For each task, the function counts the share of k-trial subsets in which every trial passed. The suite score averages those shares across tasks.

from math import comb

def pass_hat_k(trials: list[bool], k: int) -> float:
    c, n = sum(trials), len(trials)
    if k > n:
        raise ValueError("k cannot exceed the number of trials")
    return comb(c, k) / comb(n, k)  # share of k-trial subsets where all pass

results = {  # hypothetical graded trials, five per task
    "refund-over-limit-001": [True, True, False, True, True],
    "refund-under-limit-002": [True, True, True, True, True],
}
for k in (1, 3):
    score = sum(pass_hat_k(t, k) for t in results.values()) / len(results)
    print(f"pass^{k} = {score:.2f}")  # pass^1 = 0.90, pass^3 = 0.70

In this made-up run, the over-limit task failed once in five trials. Its pass^1 is 0.8, but its pass^3 is only 0.4. That single failed trial is the transcript you’d want to read before shipping.

Common mistakes

Most broken eval setups fail in one of five predictable ways.

  • Chasing public benchmarks. A leaderboard score says little about your tasks, tools, or users.
  • Running vibe checks. Eyeballing a few outputs isn’t an eval. OpenAI explicitly warns against “vibe-based evals” (OpenAI).
  • Scoring on Likert scales. A 1-5 score feels nuanced but is hard to calibrate and hard to act on.
  • Writing one-sided evals. Testing only what the agent should do trains it to overdo it.
  • Never reading transcripts. Aggregate scores hide grader bugs and strange agent behavior. Read failed transcripts as a habit, not an incident response.

Frequently asked questions

How many eval tasks does an AI agent need?

Fewer than you’d think to start. Anthropic says “20-50 simple tasks drawn from real failures is a great start” (Anthropic, 2026). Grow the suite as error analysis surfaces new failure modes. Calibrating a judge needs more: Husain and Shankar suggest 100-200 labeled examples per failure mode (Husain & Shankar, 2026).

Can I trust an LLM judge?

Only after you’ve measured it against human labels. In a 2023 chat study, GPT-4 agreed with human experts 85% of the time, versus 81% between humans (Zheng et al., 2023). The same study found position and verbosity biases. Calibrate on a held-out test set, use binary verdicts, and re-check whenever your criteria change.

What’s the difference between pass@k and pass^k?

pass@k is the chance that at least one of k attempts succeeds. pass^k is the chance that all k succeed (Yao et al., 2024). By our derived arithmetic, at 75% per-trial success and k=10, pass@k is about 99.9999% while pass^k is about 5.6%. Use pass^k for any agent that acts on a user’s behalf.

Should I use public benchmarks to evaluate my agent?

Use them to shortlist models, not to decide whether your agent works. Benchmarks saturate fast. The original SWE-bench launched with a best score of 1.96%, from Claude 2 (Jimenez et al., 2024). On its Verified subset, Anthropic notes that scores went from 40% to over 80% in one year (Anthropic, 2026). Benchmarks can also be flawed, as OpenAI’s audit showed. The eval that matters is built from your own traces.

Where to start

Agent evals aren’t polish you add after launch. They’re the spec, the test suite, and the release gate in one artifact. Start from real failures, grade outcomes cheaply, report pass^k, calibrate your judges, and test your harness like any other code.

None of this requires buying a new platform. It requires traces, one expert who owns “good,” and a few days of careful reading. If you’d like a second set of eyes on an existing agent’s evals, Moonlight AI helps teams audit and fix them.

Here’s your step for this week: pull 30 recent traces, read every one, and write a one-sentence label for each failure you find. Those labels are the start of your first eval suite. Evals tell you whether an agent is reliable; for the design controls that make it reliable, read why reliability takes design.