Skip to main content Announcing Tool Gateway MCP: the universal MCPRead the announcement
Guillaume Lebedel Guillaume Lebedel · · 6 min read
Two interchangeable model adapters connect to a shared tool layer, task records and evaluation fixtures.

Build agent architecture for routine model swaps

Table of Contents

A model upgrade should be a bounded change with a repeatable evaluation. Between 9 July and 22 September 2026, OpenAI’s Sol tier moved through three price points and a generation change. The integrations and business rules around an agent have a longer useful life than any one model choice.

That is where I would invest: reusable task fixtures, independent tool execution and task-level accounting. They reduce the work needed to test a release without assuming the new model will behave like the old one.

How quickly did model prices change?

The published price record gives a concrete reason to make migrations routine.

DateChangeInput / output per 1M tokens
9 July 2026GPT-5.6 Sol launched$5 / $30
9 July 2026GPT-5.6 Luna launched$1 / $6
24 July 2026Claude Opus 5 launched$5 / $25
30 July 2026OpenAI cut Luna pricing$0.20 / $1.20
21 August 2026OpenAI cut Sol pricing$4 / $20
22 September 2026Claude Opus 5.5 launched$4 / $20
22 September 2026GPT-6 Sol launched$2 / $10
22 September 2026GPT-6 Luna launched$0.10 / $0.50

Sources: OpenAI’s GPT-5.6 announcement and dated pricing updates, its GPT-6 Sol and Luna announcement, and Anthropic’s Opus 5 and Opus 5.5 announcements. These are listed uncached input/output rates, not complete workload bills.

Sol had an initial price and two subsequent changes. The August change was a repricing; September brought a different model. Your architecture needs to distinguish those events because only the latter requires a behavioral migration.

Public results help identify candidates. They cannot establish that your agent will improve. Artificial Analysis found lower per-task costs for GPT-6 Sol and Luna alongside regressions on its knowledge-work evaluation. A report-generation agent must test presentation quality and rubric coverage directly.

What needs checking during a model swap?

Prompts and context handling. Inspect the candidate’s actual traces. Does it gather the necessary context, produce a usable answer and stop? Longer output can exceed downstream limits; shorter output can omit a required part of the deliverable. Keep the initial prompt fixed for the first comparison, then label any later prompt tuning as a separate experiment.

Effort settings. Treat model and effort as a configuration pair. Identical setting names do not imply equal cost or quality across versions. Compare the configurations that meet the same task requirements, and keep production settings in the evaluation record.

Tool behavior and caching. Recheck argument validity, action order, permissions and repeated calls. For cache preservation, follow provider-specific instructions. OpenAI’s GPT-6 guidance keeps definitions, schemas and ordering stable while using allowed_tools or tool_choice to change availability. Effort changes use an appended configuration_update with unchanged request-level effort. Editing arbitrary schemas can still break reuse.

Output contracts. Exercise parsers against structured results, refusals, partial answers and timeouts. An agent that abstains more often can reduce incorrect answers while increasing the human queue. A parser receiving valid JSON is not evidence that the business task succeeded.

The evaluation itself. Preserve task IDs, tool fixtures and acceptance criteria. Count failed attempts in model spend and report success and fallback work separately. The re-benchmarking runbook covers paired trials and uncertainty, including when the evidence is inconclusive.

What does a concrete migration look like?

There is published evidence for comparing upgrades inside an existing agent: in Anthropic’s launch report, GitHub describes testing Opus 5.5 in Copilot CLI and VS Code, with more terminal tasks solved in fewer than half the steps in VS Code. That is a vendor-published customer result. It does not tell us GitHub’s migration time or internal architecture.

Consider a separate, illustrative support agent that looks up a customer, reads open tickets and drafts a response. This example describes the boundaries to preserve, not a measured StackOne migration.

BoundaryStable contractWhat the candidate can change
Task fixtureSame customer request and frozen ticket dataModel-generated plan and answer
Model adapterNormalized tool calls and usage recordsProvider request format, model ID and effort
Tool executionAuthorized account, action schema and validationWhich allowed actions the model requests
Acceptance checkCorrect facts, permitted actions, usable draftWhether the output passes
Routing policyTask-class rule and rollback targetCandidate configuration and rollout share

First replay the frozen requests through both adapters. Execute reads against the same snapshot and simulate writes, so one run cannot change the next run’s world. Compare the final drafts with a rubric that does not identify the model.

Suppose a trace shows the candidate repeating a ticket lookup. The tool still enforces the same permissions, and the task record exposes the extra call. Inspect the repeated lookup before changing the integration. If instead the adapter emits malformed arguments, fix that boundary and rerun the affected fixtures.

Once the candidate passes the quality, cost and latency gates, route a small share of eligible traffic to it and retain the previous configuration for rollback. The integrations remain stable while model-specific behavior is tested and deployed.

Which investments make the next swap easier?

A reusable evaluation set. Store representative tasks, known difficult cases and explicit success criteria. Keep a held-out set for the final decision so repeated prompt tuning does not teach the agent only the evaluation fixtures.

Independent tool execution. Keep authorization, input validation and action execution outside provider-specific prompt code. At StackOne, the integration layer supplies this stable boundary; tool discovery controls which actions the agent sees. A new model still needs its tool-calling behavior evaluated.

Normalized accounting. Log task ID, attempt, model version, effort, outcome and spend by billable token category. Preserve raw usage alongside normalized fields so billing changes can be checked. Different cache rates are one reason a headline price cut does not translate directly into a token reduction, as the cost-per-task analysis explains.

Configuration owned in one place. Keep model choice, effort, limits and fallback routing together. Make the selected configuration visible in traces and preserve a known working version. This makes rollback a configuration change with a clear owner.

I would also track elapsed time from a release worth evaluating to a documented decision. A slow evaluation cycle delays beneficial migrations. Reusable fixtures and explicit boundaries shorten that cycle while keeping the quality decision grounded in the work your agent performs.

Prices, benchmark observations and caching guidance checked against the linked sources on 25 September 2026.

Frequently Asked Questions

How often has the model layer repriced in 2026?
OpenAI's Sol tier had three price points and two price changes between 9 July and 22 September 2026. It launched at $5/$30 per million input/output tokens, moved to $4/$20, then GPT-6 Sol launched at $2/$10. The last change also introduced a new model generation, requiring workload evaluation.
Why can't you choose a model from a leaderboard alone?
A leaderboard measures a particular task distribution and evaluation setup. Your agent may use different tools, effort settings, prompts and acceptance criteria. Use public results to shortlist models, then compare them on representative tasks with the same success rubric, recording failures, latency and total cost across all attempts.
Is a newer, cheaper model always better for your workload?
No. A new model can reduce cost while regressing on a task category your product depends on. Artificial Analysis reported lower costs for GPT-6 Sol and Luna alongside knowledge-work regressions. Evaluate your actual deliverables and downstream parsers before treating a lower API rate as a safe migration.
What engineering work does a model swap involve?
Check prompt behavior, effort configuration, tool selection, output contracts and evaluation results. Provider adapters may also need changes to request formats and usage accounting. Reusable fixtures and independent tool execution reduce repeated work, but they do not remove the need to validate the new model's behavior on your workload.
Which parts of an agent stack can survive a model swap?
Keep task fixtures, success rubrics, tool authorization and execution, normalized usage records, and routing policy independent of model-specific request code. A provider adapter handles the changing API boundary. These separations let you compare models against stable application contracts while preserving the ability to roll back a deployment.

Put your AI agents to work

All the tools you need to build and scale AI agent integrations, with best-in-class connectivity, execution, and security.