Open-Weight Models Are Gaining Ground in Enterprise AI
For a long time, the open-weight model discussion centered mostly on developers. Downloads, experiments, local inference, fine-tunes and benchmarks gave us plenty of evidence that developers were interested, but it was harder to assess whether that interest would turn into meaningful production use. We have now begun to get some strong evidence.
According to Vercel’s AI Gateway, open-weight models accounted for about 11% of token volume in April 2026. By June, that had risen to 29%. In June, roughly one in eight enterprise AI Gateway customers was already running at least one open-weight model in production.
More recent data suggests that growth has continued. Open weights accounted for 28.4% of Vercel AI Gateway token volume on June 24. By August 22, their share had reached a record 62%. Over the latest two-month period, they accounted for about half of all token volume.
Source: Vercel
This is just one gateway and should not be treated as overall market share. But the speed of the increase is worth paying attention to, especially as the same models are also seeing adoption across other inference providers and developer platforms.
That leads to my first hypothesis:
Open-weight models are moving beyond developer experimentation and becoming a meaningful part of production AI workloads.
The Vercel data is only one view of the market, but the combination of rising token share and early enterprise usage makes the production signal difficult to ignore.
The economics behind that adoption are also worth looking at. In June, open-weight models generated 29% of Vercel’s token volume while accounting for less than 4% of spend. Anthropic, by comparison, accounted for 32% of token volume but 61% of spend.
Source: Vercel
That leads to my second hypothesis:
Open-weight models may win significant volume well before they win significant revenue.
Closed frontier models still capture a disproportionate amount of spending. Enterprise customers appear willing to pay for stronger models when the quality of the answer or the cost of failure justifies it. But they do not need the most expensive model for every task.
As enterprise customers deploy more AI, they will have an incentive to match models to workloads. The hardest reasoning task might go to a frontier model, but routine classification, extraction, repetitive agent steps, coding tasks or highly specialized workloads might go somewhere else. This also changes how I think about enterprise adoption.
There is already evidence that larger AI users are becoming heavily multi-model. Vercel found that teams processing 100 to 1,000 requests per month regularly used an average of one model. At 1,000 to 10,000 requests, that increased to three. At 100,000 to 1 million requests, it was eight. Teams processing more than 10 million requests per month regularly used an average of 35 models.
Source: Vercel
My third hypothesis:
Enterprise AI will become a multi-model market, and open weights will enter through that door.
A customer does not need to decide between Claude and Qwen, or between GPT and DeepSeek. It can use different models for different tasks. Claude or another frontier model might handle complex reasoning, whereas an open model might execute repetitive agent steps. A smaller specialized model could classify documents while another model could be tuned for a specific domain or company data. That also means enterprise adoption of open weights does not necessarily mean customers will download a model and start operating GPU clusters themselves.
Together AI says its API traffic grew from roughly 30 billion to more than 400 trillion tokens per month in nine months. Fireworks says it serves more than 40 trillion tokens per day. Fireworks also says more than 95% of its traffic comes from models specialized on customers’ proprietary data. These companies are making it possible for customers to choose and customize models without taking on all of the infrastructure work themselves.
My fourth hypothesis:
Enterprise customers will adopt open-weight models faster than they adopt self-hosted inference.
An enterprise customer can choose Qwen, DeepSeek or another open-weight model while letting an inference provider handle GPUs, scaling, reliability and much of the operational complexity. That makes open weights easier to adopt without requiring companies to build and operate their own inference infrastructure.
The distinction between “closed model through an API” and “open model running in my data center” is becoming less useful. Open-weight models can increasingly be consumed as managed services while still giving customers more choice over the model itself.
There is also a supply-side signal worth watching: what developers choose to build on. Qwen is a good example. Hugging Face recently counted more than 151,000 downstream Qwen derivatives, with developers creating roughly 180 to 210 new Qwen-derived repositories per day during the first seven months of 2026. These include fine-tunes, adapters, quantizations and other specialized versions. Downloading a model can indicate experimentation, but building on top of it represents a deeper commitment.
My fifth hypothesis:
The most important open model may not be the one that tops the benchmarks, but the one developers build on the most.
Benchmark performance still matters and is worth paying attention to. But a model family becomes more useful as developers create specialized versions, tooling, deployment options and expertise around it. Over time, that surrounding activity can become an important indicator of adoption alongside benchmark scores. Qwen also points to a broader development in the open-weight market.
Chinese labs are setting much of the pace in open-weight models.
Qwen’s developer adoption is one indication, but the trend is broader. DeepSeek, GLM, Kimi and other Chinese model families are competing near the frontier while continuing to make weights available. Chinese labs have also been aggressive about releasing large models under relatively permissive licenses.
The usage data is beginning to reflect this as well. Models from DeepSeek, Qwen and GLM are attracting meaningful production traffic across gateways such as Vercel and OpenRouter. This makes China an important part of the open-weight discussion, not only as a source of new models but also as a source of models that developers and customers are actually choosing for production.
That does not mean Chinese model companies will capture all of the economic value. A model developed in China may be served by a US hyperscaler or inference provider, run on NVIDIA or AMD hardware, routed through a gateway, and customized by another company for an enterprise workload. Open weights can distribute value across several layers: the model developer, inference provider, cloud provider, hardware vendor, gateway, and the company specializing the model for a particular workload.
There is also a demand-side factor that could push adoption further: agents.
Agentic AI could become one of the strongest economic drivers for open-weight adoption.
Agents can make many model calls to complete a single task, and that changes the economics quickly. While a chatbot making one expensive model call may be reasonable, an agent making 20 or 50 calls to complete a business process creates a much stronger incentive to manage inference costs. This could favor open-weight models in workloads where they are capable enough and considerably less expensive. As agentic workloads grow, the cost difference between models is multiplied across many calls rather than showing up once.
That does not mean frontier models lose their place. Customers will continue to pay for them when stronger reasoning, reliability or accuracy justifies the cost. But agents make the economics of model selection much harder to ignore. This is why I am less interested in whether open models will “beat” closed models and more interested in whether open weights are becoming a durable part of enterprise AI.
There are several numbers worth watching over the next year: the share of production tokens going to open weights, the percentage of enterprise customers using at least one open-weight model, the rate at which developers build on major model families, the number of inference providers supporting them, and the types of workloads customers are willing to move to open models.
Developers adopted open weights first. Then more developers started building on them. Inference providers made them easier to consume without operating the underlying infrastructure. Now we are seeing production token volume rise, providing an early signal of enterprise adoption.
The next question is how broadly enterprise customers will use open-weight models, and where the economic value will accrue as they do.