The AI industry spent the first half of 2026 arguing about which model was smartest. Its largest customers have been asking a narrower question, the one that shows up on their invoices: what does a unit of that intelligence cost?
OpenAI answered part of it on Thursday, cutting the price of GPT-5.6 Luna, the fastest and lowest-cost model in its newest family, by roughly 80%. Luna now costs $0.20 per million input tokens and $1.20 per million output tokens. The mid-range model, GPT-5.6 Terra, falls 20%, to $2 and $12. The flagship, GPT-5.6 Sol, is unchanged at $5 and $30. The cuts come three weeks after the family became generally available on July 9.
The speed is what makes this unusual. Ina Fried, who reported the reductions for Axios, pointed out that discounts like these normally arrive months after a model launches rather than weeks.
Tokens are the billing unit for AI models, each one roughly three quarters of a word. Input tokens are what you send; output tokens are what the model writes back. A single automated coding task can push millions of tokens through an interface in an afternoon. That is the pressure behind the past few months of pricing news, as OpenAI and Anthropic have both reworked rates, rate limits and usage policies while newer reasoning models consume far more tokens on long-running tasks than their predecessors did.
OpenAI says its own model helped cut the cost of running it
On Wednesday, the day before the price change, OpenAI published an engineering post explaining where the savings came from. Until now, the cost of serving a large model fell mainly through two levers: better hardware, and human engineers hand-tuning the software that runs on it. OpenAI's claim is that a third lever has opened.
The company writes that, working inside its Codex coding tool, "GPT-5.6 Sol autonomously rewrote and optimized our production kernels." Kernels are the low-level code that executes a model's mathematical operations on a graphics chip. OpenAI says that work, combined with related changes, reduced end-to-end serving costs by 20%.
For a reader outside the field, the significance is not the engineering. It is the loop. If a frontier model can meaningfully lower the cost of operating frontier models, the price curve steepens on its own, without waiting for the next chip generation. That is a different economic story from the one told by $700 billion in announced 2026 infrastructure spending.
Luna is cheaper, but it is not the cheapest
Even after an 80% cut, Luna does not undercut the market floor. DeepSeek V4 Flash, an open-weight Chinese model, runs at $0.14 input and $0.28 output. Google's Gemini 2.5 Flash-Lite sits at $0.10 and $0.40. Luna's output price remains more than 4x DeepSeek's, and output is where most agent workloads spend money.
Terra's new $2 and $12 lands exactly on Gemini 3.1 Pro's published rates, which is unlikely to be coincidence. Anthropic, meanwhile, is running Claude Sonnet 5 at a promotional $2 and $10 through August 31, reverting to $3 and $15 on September 1. The cheap and mid tiers are converging on the same numbers. The flagship tier is not.
The figures OpenAI has not published
The 20% serving-cost improvement is OpenAI's own production measurement. No outside party has audited it, and the company has not disclosed gross margin on inference at any tier, so there is no public way to tell whether Luna is profitable at $0.20 or is being priced to hold share against cheaper open-weight rivals.
Several details also work against the headline number. All three GPT-5.6 tiers carry a long-context meter that roughly doubles rates once a request exceeds the standard window. Writing new content into the prompt cache costs 1.25x the uncached input price, even though reading from it earns a 90% discount. And Sol, the model most likely to be handling the expensive multi-hour agent runs that drove the pricing complaints in the first place, got nothing.
What OpenAI has demonstrated is a floor that keeps dropping and a ceiling that does not move. Cheap intelligence is getting cheaper on a schedule now measured in weeks. Frontier intelligence still costs $30 per million words it writes, and the company just declined the chance to change that.
