// blog

Opus 5.5, Sonnet 5.5, and GPT-6 make intelligence cheaper to use

Anthropic and OpenAI are lowering the cost of useful work. Token prices explain part of it; efficiency, model routing, and verification explain the rest.

7 min read
  • ai-platforms
  • models

The most interesting question about this week's model launches is which work becomes worth automating now.

Anthropic released Claude Opus 5.5 on September 22 and Claude Sonnet 5.5 on September 28. OpenAI introduced GPT-6 Sol and GPT-6 Luna on September 22. Across both companies, the direction is similar: bring stronger capabilities into a price range where developers can use them more often.

My reading is that this changes the economics of building AI products. A model that is slightly better on a leaderboard may win a demo. A model that clears your quality bar at a fraction of the cost can change what your product does for every customer, every day.

Four launches, different ways to save

These are the published standard API rates as of September 29, 2026, in US dollars per million tokens. They exclude caching, batch discounts, premium service tiers, tool charges, and negotiated contracts. Sources are the launch announcements linked above.

Model Input Output
Claude Opus 5.5 $4.00 $20.00
Claude Sonnet 5.5 $2.00 $10.00
GPT-6 Sol $2.00 $10.00
GPT-6 Luna $0.10 $0.50

Opus 5.5 combines lower rates with greater efficiency. Its input and output prices fall 20% versus Opus 5, while cache reads fall from $0.50 to $0.20 per million tokens. Anthropic reports roughly 40% lower cost on typical workloads at default settings, reflecting both pricing and reduced token use. That 40% is a vendor estimate for a workload mix, not a universal discount.

Sonnet 5.5 is the clearest example of intelligence getting cheaper without a token-price cut. It keeps Sonnet 5's $2 input and $10 output rates. Anthropic says it uses fewer tokens and costs up to 30% less per task in its testing. The company positions it for well-scoped work, with Opus remaining stronger on complex, open-ended tasks requiring sustained judgment.

GPT-6 Sol and Luna make the price change explicit. Sol halves the published GPT-5.6 promotional rates from $4/$20 to $2/$10. Luna moves from $0.20/$1.20 to $0.10/$0.50: a 50% input reduction and approximately 58% output reduction, calculated from OpenAI's table. OpenAI attributes lower serving costs to inference and caching improvements.

Those distinctions matter. Procurement sees the rate card. Engineering sees how much work the model needs to finish the job.

Measure the cost of an accepted result

I would track the economics of an AI workflow with one primary metric:

Cost per accepted outcome = total workflow cost / outcomes that pass the acceptance criteria.

The numerator includes model calls, failed attempts, tools, execution infrastructure, and human review or repair. The denominator includes results that actually meet the product's quality bar. A response that looks finished but needs a person to redo the work is an expensive intermediate artifact.

This makes the mechanisms behind these launches easier to reason about.

First, lower token prices reduce the cost of the same workload. Second, a more efficient model can need less context, fewer reasoning tokens, or fewer tool iterations. Third, better task performance can reduce retries and escalations. Only the first benefit is visible in a price table. The others need to be measured in your own system.

Caching deserves separate attention because agents repeatedly revisit instructions, documents, and tool context. OpenAI's launch announcement describes improved cache hit rates and a 90% discount on cached input reads. That discount applies to eligible reused input, not to the whole bill. Fresh context and output still cost money.

A cheaper model can also be more expensive overall. If it makes enough mistakes, takes too many steps, or creates review work, its lower rate will not rescue the workflow. Conversely, paying more for a capable model can be economical when it resolves a difficult case cleanly.

This is why I would avoid declaring a universal winner from the launch benchmarks. Effort settings, tools, task selection, and acceptance criteria affect the result. The relevant comparison is the cheapest configuration that reliably completes your workload within its latency and risk constraints.

The small-model economics are hard to ignore

Consider an illustrative task consuming 10,000 uncached input tokens and 2,000 billed output tokens, including any billed reasoning tokens. At the rates above, its model cost is $0.08 with Opus, $0.04 with Sonnet or Sol, and $0.002 with Luna.

At one million such tasks, that becomes $80,000, $40,000, and $2,000 respectively, before tools, retries, and review. This is arithmetic under an identical token budget, not a performance benchmark or an assumption that the models are interchangeable.

The useful architectural question is how much traffic can meet the required standard at the lower price.

Suppose, purely as a scenario, 90% of tasks can be routed directly to Luna and 10% directly to Sol, using those token budgets. The model bill would be $5,800 per million tasks instead of $40,000 for Sol alone. That assumes correct routing upfront; it excludes the cost of a router, verification, and unsuccessful first attempts. An escalation pipeline would need to add those costs back.

Even after allowing for overhead, the potential is large enough to justify an evaluation. The next step is to measure whether the quality survives.

Why this matters to the market now

I see four consequences for the current AI ecosystem.

More products can afford continuous intelligence. A useful feature that runs once when a user clicks a button has different economics from one that checks every incoming document, reviews every change, or investigates every exception. Lower cost per successful task makes some of those background workflows viable. The opportunity is to revisit work that was previously excluded by its unit economics.

Application companies get more room to compete. A software company charging a fixed subscription can spend the savings on margin, higher usage limits, or more capable features. Small teams also need less inference budget to test a product against realistic traffic. My expectation is that some savings will reach customers through better products and more generous limits, especially where competitors can access the same models.

Model providers have to compete across several price levels. Opus, Sonnet, Sol, and Luna create different candidates for different workloads. Sonnet and Sol sharing the same headline rates does not make them equivalent: token consumption, latency, tools, reliability, and the surrounding platform can decide the effective price. Buyers have a stronger reason to compare completed work across providers instead of committing every task to one model family.

Differentiation moves toward the system around the model. When capable inference becomes easier to buy, proprietary context, good integrations, distribution, and measurable reliability become more valuable. A thin interface around a model has a harder time defending its price when customers can obtain similar raw capability elsewhere. Open-weight deployments also need to be compared on their full operating cost and control benefits, rather than assuming self-hosting wins on price alone.

These are implications of the launches, not measured market outcomes. They also do not imply that every AI budget will shrink. As I argued in Cheaper tokens will not mean smaller AI budgets, lower unit costs can expand the amount of work teams attempt. The opportunity to spend more productively is different from a promise to spend less.

What I would change in an AI platform

I would start by rerunning evaluations on real production tasks, including difficult cases and failures. Keep the acceptance criteria fixed, record total tokens and tool calls, and measure human correction time alongside task success and latency. A migration should earn its place on that evidence.

Then I would revisit model routing. Luna is a candidate to evaluate for routine volume; Sonnet and Sol are candidates for more demanding everyday work; Opus deserves evaluation on the ambiguous cases where additional judgment may repay its price. Those are starting hypotheses for a routing policy, not guarantees about any particular workload.

The router should use task characteristics and observable failure signals. Failed tests, missing evidence, invalid structured output, and unsuccessful tool actions are better reasons to escalate than the model's own confidence. For consequential actions, human review may remain part of the acceptance process.

I would also make the budget explicit: limits on attempts, tokens, elapsed time, and tool spending. Cheaper calls can hide an inefficient loop for a while. They cannot make an unbounded loop economical.

Finally, I would spend some of the savings on evaluations and better verification. A second model can help review a result, but agreement between models is not proof. Executable checks and evidence from the underlying system remain valuable wherever they are available.

The reason these releases matter is practical. They expand the set of tasks for which capable AI can deliver an acceptable result at an acceptable cost. The teams that benefit most will turn that opportunity into measured workflows, then keep improving the economics as the models change.

← All posts