Skip to main content Announcing Tool Gateway MCP: the universal MCPRead the announcement
Guillaume Lebedel Guillaume Lebedel · · 6 min read
Input, cache and output charges are added to model spend, divided by completed tasks.

Why cost per task replaced price per token

Table of Contents

On 22 September 2026, OpenAI announced lower API rates for GPT-6 Sol and Luna, while Anthropic said Claude Opus 5.5 costs 40% less to run than Opus 5 at default settings. These headlines measure different things. An API rate tells you what a category of tokens costs. Running cost also depends on how many tokens the agent uses, which categories they fall into and whether the task succeeds.

For a team paying an agent bill, dollars per completed task is the useful comparison. It connects the invoice to work delivered.

What is cost per completed task?

Start with each model call. Multiply each disjoint billable token category by its applicable rate, then add the charges. Sum every call across successful, failed and abandoned attempts, and divide by the number of tasks that meet your predefined success criterion.

call spend = sum(category tokens × category rate)
cost per completed task = total model spend / completed tasks

If rates are quoted per million tokens, divide each token charge by one million. Normalize the provider’s usage fields first: a total input count may already include cached input, so adding both without subtracting the overlap double-counts it. Cache writes can have different rates by retention period. Keep external tool charges and human fallback costs alongside model spend if they are outside this definition.

A rate-only change on an unchanged model can lower spend without requiring a behavioral migration. Switching to a new model or version requires evaluation even when the vendor attributes its lower price to serving improvements. GPT-6 is a model change, and its workload results differ from GPT-5.6.

Why did a 50% price cut produce different per-task savings?

Artificial Analysis’s evaluation, published on 22 September, measured more output tokens per Intelligence Index task from both new models at maximum effort.

Output tokens per Intelligence Index task, max effort
GPT-5.6 Sol
29k
GPT-6 Sol
31k
GPT-5.6 Luna
41k
GPT-6 Luna
51k

Its measured benchmark cost per task fell from $1.99 to $1.06 for Sol, about 47%, and from $0.18 to $0.07 for Luna, about 61%. These rounded benchmark figures include more than output tokens, so the chart alone cannot reconstruct the bill.

Artificial Analysis reports weighted cost per benchmark task, using task count as the denominator. The cost per completed task defined here divides all spend by successful outcomes.

Model transitionInput, per 1M tokensOutput, per 1M tokens
GPT-5.6 Sol to GPT-6 Sol$4 to $2$20 to $10
GPT-5.6 Luna to GPT-6 Luna$0.20 to $0.10$1.20 to $0.50

Source: OpenAI’s GPT-6 Sol and Luna announcement, 22 September 2026. OpenAI describes the cuts as 50% against GPT-5.6 promotional pricing; Luna’s listed output rate falls about 58%.

OpenAI attributes lower rates to caching and inference improvements. That makes the new models worth testing, but lower serving costs do not establish equivalent behavior. Artificial Analysis also measured regressions on its knowledge-work evaluation. A team producing reports or presentations needs to check deliverable quality, even if an aggregate intelligence score stays similar.

What does Anthropic’s 40% saving establish?

Anthropic’s Opus 5.5 announcement reports lower running costs at default settings and fewer tokens per task. Its pricing table contains different reductions by category:

Category, per 1M tokensOpus 5Opus 5.5Rate reduction
Uncached input$5$420%
Output$25$2020%
Cache reads$0.50$0.2060%
Cache writes, 5-minute TTL$6.25$520%

That mix matters. A workload dominated by cache reads receives a different effective rate cut from one dominated by output. You cannot subtract the 20% input/output cut from the 40% running-cost claim and attribute the remainder to fewer tokens.

In a hypothetical workload where every applicable rate falls 20% and the token mix stays constant, a 40% bill reduction would imply 25% fewer tokens: 0.60 / 0.80 = 0.75. The announcement does not establish those assumptions, so that calculation is not a measurement of Opus 5.5’s token reduction.

Effort changes the comparison too. Artificial Analysis measured roughly 119k output tokens per task for Opus 5.5 at max effort against 73k for Opus 5, with benchmark cost per task roughly level. A 20% output-rate cut alone cannot offset approximately 1.6 times the output: 0.8 × 1.6 = 1.28. The full bill also reflects input, caching and the evaluation’s usage mix.

How should you compare effort settings?

Evaluate each model at settings that could meet your product’s quality and latency requirements. A setting named “max” is not a standardized compute budget across versions or vendors. Copying the label from one model to another does not hold the experiment constant.

Record the exact model, effort, prompt, tool surface and retry policy for every run. Compare the cheapest configuration that clears the same acceptance criteria for each task class. A document extraction step may have a different acceptable configuration from a repository-wide code change.

If your evaluation uses one setting and production uses another, rerun the comparison with the production configuration before budgeting against it.

What belongs in the cost dashboard?

Lead with total model spend divided by completed tasks, then show success rate and human fallback work beside it. A falling ratio can still conceal a worse product if cheaper rates compensate for more failures.

Use these trace measurements to locate the cause of a change:

  • Preserve retry and abandoned-attempt spend under the original task ID.
  • Record which tool actions were needed, repeated or unsuccessful. Several tools can run between model calls, so tool-call count is not a direct token multiplier.
  • Separate uncached input, cache reads, cache writes and output using the provider’s accounting rules.
  • Measure the tool definitions and returned context the model receives. A growing catalogue can add prompt tokens without adding useful work; see the agent tool count ceiling.

For a trace-level breakdown, see where your agent’s tokens go.

Run representative tasks through the candidate configurations, retain failures in the totals and compare the results against your quality requirements. The model-swap evaluation runbook gives a worked example of that calculation and the uncertainty around a migration decision.

Prices and benchmark figures checked against the linked sources on 25 September 2026.

Frequently Asked Questions

What is cost per completed task?
Cost per completed task is total model spend across all attempts divided by tasks that meet your success criterion. Calculate each call using its uncached input, cache reads, cache writes and output at their applicable rates. Include failed attempts and retries; report external tool charges and human fallback costs separately.
Why did OpenAI's 50% price cut not halve every bill?
A price cut does not fix token usage or the mix of billable categories. Artificial Analysis measured more output tokens from GPT-6 Sol and Luna than their predecessors. Its benchmark cost per task fell about 47% for Sol and 61% for Luna. Those benchmark results do not predict every production workload.
How does Anthropic's 40% cost reduction differ from OpenAI's price cut?
Anthropic's 40% is a measured running-cost claim at default settings, combining pricing and usage changes. Opus 5.5 uncached input and output rates fell 20%, while cache reads fell 60%. OpenAI's headline describes API rates. Neither announcement alone determines the saving on your workload or isolates how much token usage changed.
Does the effort setting change cost per task?
Yes. Effort can change token consumption, latency and task quality, so record it alongside the model version. Anthropic's 40% saving describes default settings. Artificial Analysis tested Opus 5.5 at maximum effort and found benchmark cost per task roughly level with Opus 5 despite higher output-token usage. Compare settings you could actually deploy.
What should a platform team measure alongside model prices?
Measure failure-inclusive dollars per completed task, task success rate, latency and human fallback work on representative traffic. Break model spend into billable token categories, and record model version and effort. Tool calls, context size and retries help explain changes, but their counts cannot replace the actual token usage records.

Put your AI agents to work

All the tools you need to build and scale AI agent integrations, with best-in-class connectivity, execution, and security.