The most important LLM benchmark today: Price
The most valuable benchmark for large language models—open and proprietary—is quickly becoming price. Google launched Gemini 3.7 Flash, which is billed as its most intelligent workforce model, and included a promotional price that’ll run to January 1, 2027.
Google isn’t alone. In multiple releases for models, price is taking center stage. These model providers have read the room. Token costs matter and sticker shock has quickly turned to falling prices. Efficiency also matters because you can’t have an LLM that overthinks a task and blows the budget. All of this price compression is happening just as Anthropic and OpenAI, which are funding (via capital raises) much of the AI ecosystem.
What’s good for enterprise AI customers may not be great for the IPOs of the two giant foundation model players.
Google released Gemini 3.7 Flash, its flagship workhorse model. Gemini 3.7 Flash is designed for coding, knowledge work and agents. The company outlined the usual benchmarks, but the headliner is price. Gemini 3.7 Flash is priced at 75 cents per 1 million input tokens and $3.75 per 1 million output tokens.
We could go into other benchmarks, but who cares? Gemini 3.7 Flash will be good enough. This price theme for LLMs is popping up everywhere in earnings conference calls. There isn’t an enterprise that’s not watching the AI budget closely and looking to optimize.
Here’s a quick tour that highlights how token costs are becoming the primary selling point.
- Writer, which specializes in marketing and revenue AI workflows, released its Palmyra X6 flagship model and upgrades to its Writer Agent harness. The announcement touted Palmyra X6’s cost within Writer’s harness. The company said Writer Agent operates at a 52% lower cost.
- Meta’s move back into open weight models kicked off with Muse Glimmer. Pricing is a weapon for Muse’s comeback plan. See: Meta releases open weight Muse Glimmer model with open Muse Spark 1.2 on tap
- Nvidia’s launched its latest Nemotron model. The Nemotron family of models is essentially set up for enterprises and SaaS vendors to customize.
- SpaceXAI is also busy launching LLMs with Grok 4.6 and aggressive pricing. See: SpaceXAI, Meta puts pricing squeeze on Anthropic, OpenAI
- DeepSeek launched V4-Pro its most advanced model that’s priced at 43 cents per 1 million input and 87 cents per 1 million output tokens. DeepSeek also added dynamic pricing, which equates to a price increase to some degree, but still undercuts other models.
We could go on. Ramp data highlighted how open source AI adoption by enterprises is picking up. Everyone in the AI ecosystem is on the pricing case. Databricks announced its latest annual revenue run rate and CEO Ali Ghodsi said in a statement: “Enterprises don't just want AI that talks. They want agents working across their business that remember context, deliver accurate answers, and execute work without blowing through their budgets.”
He’s not wrong.
The AI pricing revolt is well underway and companies like Uber are getting the token budget in line. Vendors are already seeing a token budget hangover. The grown-ups are managing token costs closely and that means price matters a lot more than some obscure benchmarks that serves no business purpose. “This token model may be working for the labs, but it is not working for anyone else. It's breaking corporate budgets without results to justify the expense,” said Palantir CEO Alex Karp.
Moody’s CFO Noemie Heuland said on the company’s second quarter earnings call:
“Our internal AI and token cost today is actively governed. We have a variety of tools that we put at the disposals of our engineers, our back-office teams. We have very strict monitoring and training to ensure they are using the best tools for the task at hand. And I'm pretty proud of what we've implemented if I listened to some of my peers and all the noise around token cost explosion.”
Synchrony Financial CFO Brian Wenzel said on the company’s second quarter earnings call:
“If you go into our technology group, there's most certainly going to be a different way in which we look at the engineers and how they deploy the tools and token utilization there than you'd say someone that sits in a support function like finance or human resources. I think we have to think about how we look at that in certain processes. You may say, ‘hey, I'm going to use AI to run a process today.’ But you understand what that cost is when you think about the human capital and then any other direct dollars. We're going to have to figure out if you do AI, what's that token consumption and what's the cost of that process moving forward? That's something that's going to develop.”
Block’s Amrita Ahuja, COO, CFO & Treasurer, outlined the company’s approach to token budgets on its earnings call:
“The strategy from a cost perspective for us starts with intelligently routing our workloads, being efficient in how we think about compute and leveraging multiple models, including open source models where appropriate. And then, obviously, the technology is continuously advancing. We evolve our strategy as we see those advancements as rapidly as week-to-week or month-to-month. We don't think the right answer is to constrain developer velocity or productivity using these tools. We think the answer, as I noted, is really just to be thoughtful and intentional about how we deploy the tools.”