All posts

Right-Size AI Spend: The Cheapest Model That Passes Your Evals

Most AI work doesn't need a frontier model. Find each workload's threshold with evals, route to the cheapest model that passes, and know when not to.

Misha Druzhinin
Many railway tracks cross and branch at switches in a foggy, dark blue rail yard as a train approaches in the distance.

In May 2026, an internal OpenAI model disproved a long-standing conjecture in Erdős’s unit-distance problem, and outside mathematicians checked the result. In February, Anthropic reported that Claude Opus 4.6 had found and validated more than 500 high-severity vulnerabilities, some in code fuzzed for years. Your refund-policy agent needs none of that.

Jaya Gupta, Avanika Narayan, and Jon Saad-Falcon make this case in their X article “Right-Sizing Your Intelligence Spend”. Their argument: most enterprise work has an intelligence threshold. Once a model clears it, more intelligence buys cost and variance, not better outcomes.

We agree, and this post is our take on it. We’d push it one step further. The threshold isn’t a hunch about which tasks feel easy. It’s a measurement, and the instrument is your eval suite. Below we cover how to take that measurement, what it costs, and when right-sizing is the wrong move. We also checked the article’s numbers; four needed correcting.

One disclosure the article doesn’t make: Gupta is a partner at Foundation Capital. The firm is a listed investor in PlayerZero and led Maximor’s seed round. The article names both as examples. It does present Minions and its intelligence-per-joule results as the co-authors’ own work. Moonlight AI has no investment in, partnership with, or compensation from any company named here.

Key Takeaways

  • Most enterprise AI work has an intelligence threshold that today’s models already clear. Past that point, a bigger model or more reasoning adds cost and variance, not accuracy.
  • The threshold is an eval result. For each workload, it’s the cheapest configuration that passes your eval on real traces, across repeated trials.
  • Climb a ladder, cheapest first: no model call, small or open-weight model, mid-size model, frontier on low effort, frontier on high effort. Ship the first rung that passes.
  • Report cost per verified outcome, not tokens consumed. Token growth measures usage, not value.
  • Don’t right-size discovery work, low-volume work, or anything you can’t evaluate yet.

Most enterprise work isn’t discovery

The article’s core distinction is between two kinds of work. Discovery work looks for an answer nobody has: a proof, a zero-day, a drug target. Execution work applies known rules to a specific case: this claim, this patient, this shipment.

Discovery is valuable and rare. In the May 2025 BLS occupational survey, life, physical, and social science jobs were 1,473,260 of 155,495,730 US jobs. That’s 0.95%. The US employs 2,030 mathematicians and 20,430 physicists. Everyone else is mostly doing execution work.

Execution work fails differently. When a claims agent gets a case wrong, it’s often not because the model couldn’t reason hard enough. It didn’t know the policy, missed a field in the record, or called the wrong tool. Those are context and harness problems. A larger model doesn’t fix them; it just reasons longer around them.

That’s the intelligence threshold. Below it, the model can’t do the task. Near it, capability matters a lot. Above it, the outcome depends on context, tools, and process, while cost keeps climbing.

The intelligence threshold for one workload Conceptual diagram, not data. The horizontal axis is model capability and reasoning effort. A blue value curve rises steeply until a dashed threshold line, then flattens. An orange cost line keeps rising across the whole range. Left of the threshold the model can't do the task. Right of it, extra spend buys cost and variance, not value. Value of the output Cost per task Threshold for this workload Below: the model can't do the task Above: extra spend buys cost and variance, not value Model capability and reasoning effort
Conceptual diagram, not measured data. Each workload has its own threshold. Original diagram.

Past the threshold, extra reasoning is waste

Reasoning models spend tokens on easy problems whether or not it helps. Chen et al. described this “overthinking” problem in a December 2024 paper. They found “excessive computational resources are allocated for simple problems with minimal benefit.”

The source article cites Amazon researchers for “7 to 10x as many tokens as necessary.” The actual line, from an Amazon Science blog post by a product manager, is narrower. Reasoning models use “seven to 10 times as many tokens as non-reasoning models” for comparable accuracy on simple tasks. That’s still a strong point. It just compares model types, not an ideal minimum.

Vendors say the same in their docs. Anthropic’s effort parameter documentation runs from low to max. It warns that max can lead to overthinking on less demanding tasks. When the vendor tells you the top setting can hurt, believe it.

Cost isn’t the only damage. Overqualified agents add variance. To borrow the article’s example, a password-reset agent that considers twelve explanations has twelve chances to do something odd. We cover why extra machinery tends to cost reliability in Complexity Is Easy. Reliability Takes Design.

Benchmarks can’t tell you where your threshold is

Smaller models are closing the gap, but unevenly. Qwen’s model card for Qwen3.8-27B, a 27-billion-parameter open-weight model, compares it with Claude Opus 4.6 at max effort. The 27B model wins five of the seven coding and agent benchmarks below. It loses two.

A 27B open-weight model against a frontier model Dumbbell chart of seven benchmark scores from Qwen's model card. Qwen3.8-27B versus Opus 4.6 Max: SWE-bench Pro 61.7 vs 53.4; LiveCodeBench v6 90.3 vs 88.8; OSWorld-Verified 84.3 vs 72.7; AndroidWorld 81.9 vs 62.0; CoWorkBench 70.7 vs 68.2; Terminal-Bench 2.1 73.0 vs 78.2; NL2Repo 42.3 vs 47.6. Qwen leads on five, Opus on two. Opus scores are officially reported scores as listed by Qwen. A 27B open model vs. a frontier model Benchmark score, as reported on Qwen's model card Qwen3.8-27B Opus 4.6 Max Qwen Opus 40 50 60 70 80 90 SWE-bench Pro LiveCodeBench v6 OSWorld-Verified AndroidWorld CoWorkBench Terminal-Bench 2.1 NL2Repo 61.753.4 90.388.8 84.372.7 81.962.0 70.768.2 73.078.2 42.347.6 Source: Qwen3.8-27B model card, Hugging Face (2026)
Source: Qwen3.8-27B model card, 2026. Vendor-reported; Opus scores are the officially reported figures Qwen lists.

Read that chart two ways. First, a model small enough to self-host now beats a frontier model on several agent tasks. That’s the article’s point, and it holds. Second, the losses are on long-horizon terminal work and whole-repo generation. If your workload looks like those, the small model may sit below your threshold.

Neither reading tells you what happens on your claims queue. These are vendor-reported numbers on public tasks. Your threshold depends on your inputs, your tools, and your definition of correct. Only one instrument measures that.

The threshold is an eval result

Here’s where we go further than the article. “Route easy work to cheaper models” is good advice with a silent failure mode. If you can’t prove the cheap path still gets the answer right, you’ve cut quality and called it savings. Nobody notices until customers do.

So define the threshold operationally. For a given workload, it’s the cheapest configuration that passes your eval on real traces, across repeated trials, at the bar you set. Below that configuration, the eval fails. Above it, you’re paying for headroom you can’t use.

That needs three things you may not have yet:

  • A task set from production. Sample real requests, including the awkward ones, not synthetic examples. Our eight-step method for agent evals covers building one from real failures.
  • A grader you trust. Prefer code checks: the refund amount matches policy, the right tool was called, the record changed. Use an LLM judge only where code can’t decide, and calibrate it against human labels.
  • Repeated trials. One lucky pass proves little. Report pass^k, the chance all k runs succeed. A cheaper model can look fine on pass@1 and still fail on consistency.

Climb the right-sizing ladder

With an eval in hand, test configurations from cheapest to most expensive. Ship the first one that clears the bar. We use five rungs.

The right-sizing ladder Five rungs tested from cheapest to most expensive. Rung 0: no model call, using a rule, lookup, or cache. Rung 1: a small or open-weight model. Rung 2: a mid-size hosted model. Rung 3: a frontier model on low effort. Rung 4: a frontier model on high effort. If a rung fails the eval, move down to the next. Ship the first rung that passes. Cost per call rises down the ladder. Test in order. Ship the first rung that passes. 0. No model call Rule, lookup, template, or cached answer 1. Small or open-weight model Local, self-hosted, or fine-tuned for the task 2. Mid-size hosted model The vendor's fast tier, reasoning off or low 3. Frontier model, low effort Top model, small reasoning budget 4. Frontier model, high effort Reserve for work that fails everywhere else fails eval fails eval fails eval fails eval cheapest cost per call rises most expensive
The right-sizing ladder. Each workload gets its own rung, and the eval decides. Original diagram.

Rung 0 is the one teams skip. A lot of agent traffic has a fixed answer: the same five policy questions, the same lookup, the same template. If a rule or a cache passes the eval, the cheapest model call is the one you don’t make. The source article makes this point too: the best outcome may be “no model call at all.”

The search itself is a few lines of code. This sketch is framework-neutral. You supply each rung’s run function, a measured cost per call, and a grader from your eval suite.

from dataclasses import dataclass
from typing import Callable

@dataclass
class Rung:
    name: str
    run: Callable[[dict], dict]   # task input -> output
    cost_per_call: float          # measured from your traces, incl. retries and escalations

def passes_all(rung: Rung, task: dict, grade: Callable, k: int) -> bool:
    """pass^k for one task: every one of k trials must pass."""
    return all(grade(task, rung.run(task["input"])) for _ in range(k))

def cheapest_passing_rung(rungs, tasks, grade, bar=0.95, k=3):
    for rung in sorted(rungs, key=lambda r: r.cost_per_call):
        rate = sum(passes_all(rung, t, grade, k) for t in tasks) / len(tasks)
        print(f"{rung.name}: pass^{k} = {rate:.1%}")
        if rate >= bar:
            return rung
    return None  # nothing clears the bar: fix context or tools before buying intelligence

Expect two things when you run this search. First, the None branch is the most useful result. If even rung 4 fails, check context and tools before buying more intelligence. Second, a rung can pass overall and fail one segment, such as refunds over a limit. Split those into their own workload with their own rung.

Then change what you report. The article calls it cost per verified resolution; we say cost per verified outcome. It’s total spend divided by outcomes that passed verification. Total spend includes model calls, infrastructure, retries, and the human time spent on escalations. A cheap model that escalates half its cases isn’t cheap.

What the routing research shows

Routing between models isn’t a new idea. The published results point the same way: cheap models can carry much of the work, with a strong model where it counts.

Study Approach Reported result
FrugalGPT (Chen, Zaharia, Zou, 2023) Cascade: try cheaper models first, escalate on low confidence Matched GPT-4 with up to 98% cost reduction
RouteLLM (LMSYS, 2024) Learned router between a strong and a weak model Over 85% cost reduction on MT Bench, 45% on MMLU, 35% on GSM8K, at 95% of GPT-4’s quality
Minions (Narayan et al., 2025) Cloud model decomposes, small local models execute MinionS kept 97.9% of cloud accuracy at 18% of the cost; the simpler Minion protocol kept 87% at 30.4x lower cost
GPT-5 (OpenAI, 2025) Built-in real-time router between a fast and a thinking model Product feature; no cost figure published

One correction to the article here. It describes Minions’ 30.4x result as “87.9%” from “a more aggressive configuration.” In the paper, it’s 87% from a different, simpler protocol. The headline result stands: in Minions, the expensive model was most useful for deciding what to do, not for doing every step.

Two more numbers in the article need dates. Its “18x in 16 months” intelligence-per-joule figure comes from the Intelligence per Watt paper. It combines 3.0x from models and 5.9x from accelerators, for local model-and-hardware pairs, from April 2024 to August 2025. The paper’s headline is a 5.3x gain from 2023 to 2025. The article’s 320x growth in reasoning tokens comes from OpenAI’s December 2025 enterprise report. It covers API usage in 2025, not the year before publication.

Labs sell tokens; you need outcomes

The sharpest section of Gupta, Narayan, and Saad-Falcon’s article is about incentives. A model vendor’s usage dashboard rises when you think longer, retry more, and spawn more agents. Your business gains when the same outcome needs fewer of those.

The data shows the gap. OpenAI’s 2025 enterprise report says API reasoning-token use per organization grew “approximately 320x in the past 12 months.” Meanwhile, 56% of the 4,454 CEOs in PwC’s 2026 CEO Survey reported “no significant financial benefit” from AI to date. These are different samples, so one doesn’t explain the other. Side by side, they’re hard to square: usage is climbing fast, while most CEOs don’t yet report a financial return.

We’d state the incentive point more carefully than the article does. Vendors do ship cheaper paths: small model tiers, effort controls, built-in routers, and docs that warn against maximum settings. The difference is that only you can see your outcomes. No vendor knows whether your refund was correct. So the routing decision has to live with you, wired to your evals, not to a provider default.

When right-sizing is the wrong move

Right-sizing has costs of its own. Skip it, or delay it, in these cases:

  • Discovery work. Security research, hard debugging, and novel analysis are where frontier capability earns its price. Don’t route those down.
  • Low volume. A workload that runs 200 times a month won’t repay an eval suite, a router, and a second model in production. Use the strong model and move on.
  • No eval yet. Without one, routing down is a guess. Build the eval first, even a small one, and keep the expensive model until it exists.
  • Self-hosting overhead. Open weights are cheap per token, not per month. GPUs, upgrades, security patching, and on-call add up. Price the whole stack.
  • The router itself. A router is a new component that can misclassify quietly. Log its decisions, sample them, and evaluate them like any other model.
  • Moving thresholds. New model versions, new input mixes, and new policies all shift the threshold. Re-run the ladder on a schedule, not once.

Common questions about right-sizing

Should every request go through a router?

No. Start with static assignment: one rung per workload, chosen by the ladder. Add per-request routing only when a workload mixes easy and hard cases you can tell apart cheaply. A static choice is easier to test and debug.

Is a smaller model less reliable?

Not inherently. On execution work, a well-instructed small model with good context can be more consistent than a frontier model that overthinks. Measure it with pass^k on your tasks. Don’t assume either way.

Where should we spend the frontier budget?

On decomposition, hard judgment calls, and verification of high-stakes outputs. Those are the steps where capability changes the outcome. Routine steps inside a workflow usually belong on a lower rung, as we discuss in graph engineering for AI agents.

Where to start this week

Gupta, Narayan, and Saad-Falcon end on a good line: progress means intelligence consumed per successful outcome falling toward zero. We’d add that you can’t watch it fall without a measurement. Evals make the threshold visible, the ladder finds the cheapest rung, and cost per verified outcome tells you whether it worked.

If your AI bill is growing faster than your results, Moonlight AI helps teams build the evals and routing to fix that.

Your step for this week: pull your top workload by model spend. Sample 50 real traces and write a code-based grader for the one outcome that matters most. Run your current model and the next rung down, three trials each. Compare pass^3 and cost per verified outcome. If the cheaper rung clears your bar, you’ve found money. If it doesn’t, you’ve found your threshold.