Cheaper tokens will not mean smaller AI budgets
OpenAI has cut the API price of GPT-5.6 Luna by 80% and Terra by 20%. The new list prices are $0.20 per million input tokens and $1.20 per million output tokens for Luna, and $2 and $12 respectively for Terra. Sol stays where it was, with a new Fast mode offering up to 2.5 times the speed for twice the price.
Read that as a procurement announcement and the conclusion is obvious: companies using OpenAI should reduce their AI cost forecasts.
I think that is mostly wrong.
The price cut changes the cost of a unit. Big AI customers do not forecast units in isolation; they forecast products, users, automated workflows, and gross margin. When the unit gets dramatically cheaper, the first response is rarely to run the same workload and return the savings. It is to revisit all the workloads that did not make economic sense last week.
The 80% matters more than the model name
An 80% reduction is not normal cloud price erosion. It moves an application across architecture boundaries.
At Luna's old price, a team might use one structured-output call to classify a support ticket. At the new price, it can afford a loop that reads account history, calls two tools, checks policy, drafts a response, and verifies it. The output is not merely cheaper; it is a different product.
OpenAI is explicit about this. Luna can use tools and complete multi-step workflows, and the announcement positions it for high-volume document analysis, customer-interaction classification, and routine implementation. That is the territory where enterprise volumes become enormous: small amounts of intelligence applied to every transaction rather than frontier intelligence reserved for exceptional cases.
The simple revenue arithmetic shows the bet. On like-for-like traffic, Luna needs five times as much volume after an 80% price cut for model revenue to stay flat. Terra needs 25% more volume after its 20% reduction. OpenAI is betting that the demand curve is elastic enough to clear those bars — while keeping Sol's premium intact for work where capability matters more than price.
The forecast changes by customer type
"Big AI customer" hides several different businesses.
Fixed-price software products are the clearest winners. A company selling an AI assistant per seat absorbs the inference cost while revenue stays roughly fixed. If quality and usage remain constant, cheaper Luna and Terra improve gross margin immediately. More realistically, the company spends part of the saving on higher limits, background agents, or features that were previously too expensive to enable by default. Margin improves, but less than the headline 80% suggests because the product gets more ambitious.
Usage-priced platforms have a different equation. Azure and other model distributors can pass lower prices through to customers, which reduces revenue per token but makes the platform useful for more workloads. The forecast shifts from price to throughput: lower unit revenue, higher consumption, more demand for storage, databases, identity, observability, and every other cloud service the agent touches.
Large internal adopters should revisit their automation backlog. The right question is no longer "how much does our current pilot save?" It is "which processes now cross our minimum return threshold?" A price reduction can turn an AI budget from ten expensive experiments into a shared operating layer across support, engineering, finance, and security.
And negotiated enterprise contracts complicate all of this. The public API price is not necessarily the effective price paid by the largest customers, nor will every contract reprice immediately. Nobody should multiply their current token count by 20% and call it a forecast.
Microsoft benefits; the other hyperscalers get a warning
Microsoft is the big-tech company with the most direct exposure. It distributes OpenAI models through Azure and embeds them across products such as Microsoft 365 Copilot and GitHub Copilot. Lower inference costs can improve the economics of those fixed-price products or allow more agent activity inside the existing subscription.
The timing is useful. In its latest results, Microsoft reported more than 30 million paid Microsoft 365 Copilot seats and Azure growth of 43%. Its additions to property and equipment reached $115.9 billion for the fiscal year. A cheaper workhorse model helps turn that infrastructure into more sellable outcomes, but it does not make the infrastructure bill disappear.
Google, Amazon, and Meta are different. They are primarily operators of competing AI stacks, not ordinary retail buyers of OpenAI tokens. For them, this announcement is less a discount and more a price signal. OpenAI has reset what customers expect a capable small agent model to cost. Rival models now have to match the price-performance curve, differentiate on capability or integration, or accept losing commodity inference volume.
Google's own numbers show why I would not forecast an industry-wide spending slowdown. In July, the company said its APIs were processing about 22 billion tokens per minute, up from 16 billion one quarter earlier, while demand remained supply constrained. Its Q2 update reported 82% Cloud revenue growth and described continued efficiency gains alongside expanding usage. Falling inference cost and rising infrastructure demand are happening at the same time.
Cheaper inference is not cheaper ambition
OpenAI says GPT-5.6 helped optimize the systems serving it: production-kernel changes reduced end-to-end serving cost by 20%, while experiments improved token-generation efficiency by more than 15%. Passing some of that gain to customers is good engineering and aggressive competition.
It is not evidence that the AI buildout is winding down.
My forecast for large customers would change in five places:
- Cost per completed workflow falls, especially for high-volume, tool-using tasks.
- Usage forecasts rise as background and always-on agents become economical.
- Gross-margin forecasts improve temporarily for fixed-price AI products, until higher usage consumes part of the gain.
- Model routing becomes financially material: Luna handles the volume, Terra handles ambiguity, and Sol protects the difficult tail.
- Capital spending stays high because lower prices create demand for more inference rather than satisfying a fixed quantity of it.
The risk is that demand does not expand quickly enough. Five times the Luna volume is a serious hurdle, and competitors will answer. But for customers, the direction is clearer: do not book the entire reduction as savings. Treat it as newly available product budget.
In AI, cheaper tokens are rarely returned to the CFO. They are spent on a larger loop.