OpenAI cuts price of GPT-5.6 Luna, GPT-5.6 Terra
OpenAI said it is cutting the price of GPT-5.6 Luna by about 80% and the price of GPT-5.6 Terra by 20% after improving performance of the underlying systems.
The moves, outlined in a blog post, are notable for the following reasons:
- Frontier AI companies are getting the message that the idea that enterprises will endure inflated token costs is bunk.
- Open models from China as well as the US are going to apply pricing pressure to frontier labs.
- Frontier labs can use infrastructure as a moat assuming they can continually optimize.
OpenAI said:
"Our efficiency edge comes from improving the models, the inference systems that run them, and the agentic harness that connects them to tools and context. GPT‑5.6 models take a more direct path through work. Better routing keeps hardware productive, optimized production software generates tokens more efficiently, and smarter context management helps agents avoid repeating completed work. Together, these improvements let us complete more useful work with the same compute, reducing the time, tokens, and cost required for each result."
The company added that it is using GPT-5.6 Sol is helping it deliver efficiency gains going forward. OpenAI noted that Sol "autonomously rewrote and optimized production kernels, designed and ran hundreds of experiments to improve token generation, and monitored training, intervening when problems arose."
- Get ready for US open LLMs and just in time
- Rightsizing open models may cut your AI inference spend
- SpaceXAI, Meta puts pricing squeeze on Anthropic, OpenAI
- Moonshot AI launches Kimi K3
- DeepSeek's real impact happening now
- Here’s what we learned about AI projects from enterprise buyers so far
- AI inference costs are going to be a big concern: What's the fix?
OpenAI said API pricing will go like this:
- $2 per million input tokens and $12 per million output tokens for Terra.
- 20 cents per million input tokens and $1.20 per million output tokens for Luna.
- Sol pricing remains unchanged.
- ChatGPT and Codex subscription prices and quota budgets are unchanged.
CFO Sarah Friar followed up the pricing move in a blog post that noted:
These are not simply changes to a price list. They expand the range of work that becomes practical and give customers more flexibility to balance intelligence, speed, reliability, and cost.
The right question is not which model belongs to which task. It is how much intelligence the outcome demands for the required result, how quickly the intelligence is needed, and what the intelligence should cost to achieve. That balance may change several times within the same workflow.
Customers do not buy tokens for their own sake. They want the support issue resolved, the software shipped, the contract reviewed, or the scientific question answered. The right measure is the cost of a successful outcome, including the time, retries, oversight, and errors required to get there.