The most important LLM benchmark today: Price

Published August 13, 2026

The most valuable benchmark for large language models—open and proprietary—is quickly becoming price. Google launched Gemini 3.7 Flash, which is billed as its most intelligent workforce model, and included a promotional price that’ll run to January 1, 2027.

Google isn’t alone. In multiple releases for models, price is taking center stage. These model providers have read the room. Token costs matter and sticker shock has quickly turned to falling prices. Efficiency also matters because you can’t have an LLM that overthinks a task and blows the budget. All of this price compression is happening just as Anthropic and OpenAI, which are funding (via capital raises) much of the AI ecosystem.

What’s good for enterprise AI customers may not be great for the IPOs of the two giant foundation model players.

Google released Gemini 3.7 Flash, its flagship workhorse model. Gemini 3.7 Flash is designed for coding, knowledge work and agents. The company outlined the usual benchmarks, but the headliner is price. Gemini 3.7 Flash is priced at 75 cents per 1 million input tokens and $3.75 per 1 million output tokens.

Gemini 3.7 Flash

We could go into other benchmarks, but who cares? Gemini 3.7 Flash will be good enough. This price theme for LLMs is popping up everywhere in earnings conference calls. There isn’t an enterprise that’s not watching the AI budget closely and looking to optimize.

Here’s a quick tour that highlights how token costs are becoming the primary selling point.

We could go on. Ramp data highlighted how open source AI adoption by enterprises is picking up. Everyone in the AI ecosystem is on the pricing case. Databricks announced its latest annual revenue run rate and CEO Ali Ghodsi said in a statement: “Enterprises don't just want AI that talks. They want agents working across their business that remember context, deliver accurate answers, and execute work without blowing through their budgets.”

He’s not wrong.

The AI pricing revolt is well underway and companies like Uber are getting the token budget in line. Vendors are already seeing a token budget hangover. The grown-ups are managing token costs closely and that means price matters a lot more than some obscure benchmarks that serves no business purpose. “This token model may be working for the labs, but it is not working for anyone else. It's breaking corporate budgets without results to justify the expense,” said Palantir CEO Alex Karp.

Moody’s CFO Noemie Heuland said on the company’s second quarter earnings call:

“Our internal AI and token cost today is actively governed. We have a variety of tools that we put at the disposals of our engineers, our back-office teams. We have very strict monitoring and training to ensure they are using the best tools for the task at hand. And I'm pretty proud of what we've implemented if I listened to some of my peers and all the noise around token cost explosion.”

Synchrony Financial CFO Brian Wenzel said on the company’s second quarter earnings call:

“If you go into our technology group, there's most certainly going to be a different way in which we look at the engineers and how they deploy the tools and token utilization there than you'd say someone that sits in a support function like finance or human resources. I think we have to think about how we look at that in certain processes. You may say, ‘hey, I'm going to use AI to run a process today.’ But you understand what that cost is when you think about the human capital and then any other direct dollars. We're going to have to figure out if you do AI, what's that token consumption and what's the cost of that process moving forward? That's something that's going to develop.”

Block’s Amrita Ahuja, COO, CFO & Treasurer, outlined the company’s approach to token budgets on its earnings call:

“The strategy from a cost perspective for us starts with intelligently routing our workloads, being efficient in how we think about compute and leveraging multiple models, including open source models where appropriate. And then, obviously, the technology is continuously advancing. We evolve our strategy as we see those advancements as rapidly as week-to-week or month-to-month. We don't think the right answer is to constrain developer velocity or productivity using these tools. We think the answer, as I noted, is really just to be thoughtful and intentional about how we deploy the tools.”