Skip to main content Announcing Tool Gateway MCP: the universal MCPRead the announcement
Guillaume Lebedel Guillaume Lebedel · · 8 min read
The same frozen task set runs through current and candidate models before a shared quality, cost and latency decision.

Re-benchmark your agent before a model swap

Table of Contents

Before changing an agent’s model, run the same representative tasks through the current and candidate configurations. Compare total spend per completed task alongside success, latency and human fallback work. This runbook produces that comparison and makes an inconclusive result explicit.

You need saved task inputs, repeatable tool responses, access to usage records and a success rubric. The worked example uses Python 3’s standard library and illustrative records, so it can run without provider credentials or paid model calls.

1. Define completion and freeze the comparison

Write the success criterion before running either model. For a support agent, success might require a draft that uses the right account’s records, states the relevant facts and performs no unauthorized action. Valid JSON or a fluent answer is not enough.

Many production tasks require several model calls. Put a task ID and attempt ID on every call, and keep retries under the original task. Record the model version, prompt version, effort, tool definitions and retry policy.

Freeze the task inputs and relevant tool data. Use recorded reads or a resettable test environment, and simulate writes. Otherwise one run may change the state that the next run sees. Alternate or randomize execution order to reduce time-of-day and provider-load effects on latency.

Keep the first comparison’s prompt and tools fixed. If the candidate needs adaptation, label that as a second configuration and assess it on held-out tasks.

2. Choose paired tasks and repeated trials

A set of 100-300 representative tasks can be an initial screen. It is not automatically enough to resolve a one-percentage-point success difference: that difference amounts to only one to three outcomes before you split into categories.

Sample the traffic mix, including long tasks, tool failures and cases that need human review. Keep a separate set of critical cases even if they are rare. Record category weights so an oversampled difficult category does not silently change the aggregate result.

Run both configurations on each task and repeat stochastic trials. Choose the repetition budget before examining results. Score outputs with the same rubric, preferably without showing the grader which model produced them.

Estimate uncertainty on the paired difference. For example, bootstrap task IDs with replacement, carrying both models and every repeat for each selected task together. Treating repeats as unrelated tasks overstates the effective sample size. Use category-aware resampling and fixed traffic weights when the evaluation is stratified.

If the interval is too wide for your tolerance, the result is inconclusive. More repeated trials help measure model variability; more distinct representative tasks help establish how broadly the result applies.

3. Aggregate spend without dropping failures

Normalize each provider’s usage into disjoint billable categories: uncached input, cache reads, applicable cache writes and output. Apply the relevant rates, including units and retention tiers. Some providers include cached input in total input, so normalize before adding charges.

For each evaluated task run, sum every call and retry before recording the final outcome. Then calculate:

cost per completed task = spend across all task runs / successful task runs
success rate = successful task runs / all task runs

With no successful runs, cost per completion is undefined; report spend and failures directly. When you repeat tasks, each planned repeat is one evaluated task run, while internal retries remain part of that run’s spend.

Keep external tool charges and human fallback cost separate unless you explicitly include them in the cost definition. Also report p95 latency, abstention rate and human minutes. A lower model bill cannot tell you whether the support queue grew.

A runnable example

Save this as compare.py and run python3 compare.py. Each CSV row is one paired trial; each dollar amount includes all model calls and retries for that trial. These illustrative eight pairs cover four tasks with two repeats each. They demonstrate the calculation and are far too small to validate a production migration.

import csv
import io
import random
from collections import defaultdict

records = """task,trial,current_usd,candidate_usd,current_ok,candidate_ok,current_minutes,candidate_minutes
A,1,1.00,0.60,1,1,0,0
A,2,1.00,0.60,1,1,0,0
B,1,1.00,0.60,1,1,0,0
B,2,1.00,0.60,1,1,0,0
C,1,1.00,0.60,1,0,0,4
C,2,1.00,0.60,1,1,0,0
D,1,1.00,0.60,0,0,4,4
D,2,1.00,0.60,0,0,4,4
"""
rows = list(csv.DictReader(io.StringIO(records)))
for model in ("current", "candidate"):
    spend = sum(float(row[f"{model}_usd"]) for row in rows)
    successes = sum(int(row[f"{model}_ok"]) for row in rows)
    minutes = sum(int(row[f"{model}_minutes"]) for row in rows)
    cost = f"${spend / successes:.2f}" if successes else "undefined"
    print(f"{model}: {cost}/completion, "
          f"success={successes / len(rows):.1%}, fallback={minutes} min")

by_task = defaultdict(list)
for row in rows:
    by_task[row["task"]].append(
        int(row["candidate_ok"]) - int(row["current_ok"])
    )
deltas = [sum(values) / len(values) for values in by_task.values()]
rng = random.Random(0)
samples = sorted(
    sum(rng.choices(deltas, k=len(deltas))) / len(deltas)
    for _ in range(10000)
)
low = samples[int(0.025 * (len(samples) - 1))]
high = samples[int(0.975 * (len(samples) - 1))]
print(f"Success delta 95% bootstrap interval: "
      f"{low * 100:.1f} to {high * 100:.1f} pp")

Expected output:

current: $1.33/completion, success=75.0%, fallback=8 min
candidate: $0.96/completion, success=62.5%, fallback=12 min
Success delta 95% bootstrap interval: -37.5 to 0.0 pp

The candidate’s model cost per completion falls 28%, even though it completes fewer runs and adds four minutes of human work. If the predeclared quality tolerance were a one-percentage-point decline, this interval would cross the threshold of −1 point. Hold the migration: the quality result is inconclusive. Inspect task C and collect more representative evidence.

This simple bootstrap assumes equally weighted tasks with the same number of repeats. With only four tasks, its interval is a teaching example, not a reliable production confidence statement. For a real decision, also estimate uncertainty for cost and latency using the same paired task groups; resamples with zero completions need to remain visible as undefined cost ratios.

4. Inspect abstentions and effort before calling a win

An abstention can prevent a wrong answer and still create work for a person. In Artificial Analysis’s GPT-6 evaluation, Sol’s AA-Omniscience hallucination rate fell from 92% to 60%, while questions attempted fell from 99% to 83% and overall accuracy fell from 59% to 54%.

Those metrics use different denominators; a lower hallucination rate is not the same as a higher completion rate. Your success rubric and fallback records should reveal the tradeoff. The example above also shows why model cost per completion alone can miss it.

Effort is another source of misleading comparisons. Anthropic’s Opus 5.5 claim describes lower running cost at default settings, while Artificial Analysis’s maximum-effort evaluation found per-task cost roughly level with Opus 5. Record and test the settings you would deploy. Do not transfer a saving from one configuration into another’s budget.

5. Record a decision and keep the fixtures

Before the run, write down the acceptable quality loss, desired cost change, latency limit and critical failure categories. Define three outcomes:

OutcomeDecision rule
Promote to a limited rolloutEvidence clears all gates, including uncertainty, and critical cases pass
Reject the candidateA critical requirement is breached or evidence supports an unacceptable regression
InconclusiveUncertainty spans a decision threshold, category coverage is insufficient or fixtures do not represent production

Review per-category results before the aggregate. A large saving on easy tasks should not hide failures in a business-critical category. During a limited rollout, monitor the same task metrics and retain the previous model configuration for rollback.

Save the fixture set, configuration, usage records, grader output and decision together. Those reusable artifacts make the next model swap quicker to evaluate.

Release and benchmark examples checked against the linked sources on 25 September 2026.

Frequently Asked Questions

Why isn't a public benchmark enough to decide a model swap?
A public benchmark uses its own tasks, tools, prompts and scoring rules. Your product can fail in categories that an aggregate score hides. Use public results to select candidates, then run paired evaluations on representative traffic with predefined success criteria, repeated trials and uncertainty estimates before deciding whether to migrate.
What should you measure when comparing agent models?
Measure total model spend across all attempts divided by successful task runs, plus success rate, latency, abstentions and human fallback work. Record token usage by billing category and the exact configuration. Successful-task cost percentiles are useful diagnostics, but they exclude failures and cannot replace the overall cost-per-completion ratio.
What is the abstention trap in model evaluation?
A model may reduce incorrect answers by declining more tasks. That can be useful, but it may also increase human work and reduce completion. Track success, abstentions and fallback time together. A sufficiently large price cut can still improve model cost per completion while more work moves to people.
Why must a model comparison record effort settings?
Effort affects token use, latency and quality. Settings with the same label can behave differently across models, so record the exact setting and compare configurations that meet the same task requirements. A maximum-effort benchmark does not predict the cost or quality of a production agent running at a lower setting.
When is there enough evidence to swap models?
Swap when the candidate meets predeclared quality, cost and latency gates, with uncertainty narrow enough to support the decision and no unacceptable critical failures. If the interval crosses your quality tolerance, collect more evidence or keep the current model. A promising point estimate alone is not a deployment decision.

Put your AI agents to work

All the tools you need to build and scale AI agent integrations, with best-in-class connectivity, execution, and security.