Home / Research / Open-Weight Model Adoption Statistics
Market research · September 24, 2026
Open-Weight Model Adoption Statistics: 2026 Report
This report looks at how companies use open-weight models such as Meta's Llama, Alibaba's Qwen, DeepSeek, Mistral, Google's Gemma, and OpenAI's gpt-oss alongside closed APIs. Open-weight models are models whose weights are publicly released, so organizations can download them and run inference on their own hardware or through a hosting provider.
We surveyed 640 organizations with 100 or more employees in North America and Europe between March and August 2026, collecting quarterly usage figures from Q3 2024 to Q2 2026. We combined our data with published figures from Menlo Ventures, OpenRouter, Andreessen Horowitz, Hugging Face, and Vercel.
We found that 46% of organizations ran at least one open-weight model in production at the end of Q2 2026, up from 31% in Q3 2024. Open-weight models handled 24% of those organizations' LLM tokens but only 18% of their production workloads, because they are concentrated in high-volume tasks such as document extraction. Chinese-origin models accounted for 28% of open-weight tokens, while 36% of organizations formally restrict them.
Open-Weight Model Adoption by Company Size
In the table below, we show the share of organizations running open-weight models in production, along with the open-weight share of production workloads and of monthly tokens, as of the end of Q2 2026. We count an organization as a production user when at least one open-weight model serves a live customer-facing or employee-facing workload.
The Open-Weight Model Adoption by Company Size, 2026
| Company Size | Respondents | Running Open-Weight Models in Production | Open-Weight Share of Production Workloads | Open-Weight Share of Tokens |
|---|---|---|---|---|
| 100 to 999 employees | 228 | 38% | 21% | 27% |
| 1,000 to 9,999 employees | 246 | 47% | 17% | 24% |
| 10,000+ employees | 166 | 56% | 14% | 21% |
| All respondents | 640 | 46% | 18% | 24% |
Adoption rises with company size, while depth of use falls. Organizations with 10,000 or more employees were 18 percentage points more likely than those with 100 to 999 employees to run an open-weight model in production, yet open-weight models handled 14% of their workloads against 21% at the smaller firms. Large enterprises tend to add an open-weight model for one or two high-volume tasks while keeping closed APIs as the default for everything else. Smaller companies that adopt open-weight models more often build a larger share of their stack on them, frequently through a single engineering team that standardizes on one model family. The token column tells the same story from another angle. At every company size, the open-weight share of tokens exceeds the open-weight share of workloads, by 6 to 7 percentage points, which means the workloads moved to open-weight models are heavier than average. Respondents described these as batch jobs over large document sets, where per-token price dominates the purchasing decision.
Our 24% token share is well above the 11% open-source share of enterprise LLM usage that Menlo Ventures reported in December 2025. Our figure is measured half a year later and includes mid-sized companies, which Menlo's enterprise sample weights less heavily. Closed providers still dominate: the Andreessen Horowitz CIO survey of 100 Global 2000 companies, published in January 2026, found OpenAI models in production at 78% of respondents. Open-weight models have become a common second supplier inside large companies, while the primary relationship for most enterprises remains a closed API provider.
Open-Weight Model Adoption Trend, Q3 2024 to Q2 2026
The table below tracks the same three measures for our full panel at the end of each quarter over the past two years.
The Open-Weight Model Adoption Trend, Q3 2024 to Q2 2026
| Quarter | Running Open-Weight Models in Production | Open-Weight Share of Production Workloads | Open-Weight Share of Tokens |
|---|---|---|---|
| Q3 2024 | 31% | 14% | 13% |
| Q4 2024 | 33% | 15% | 14% |
| Q1 2025 | 35% | 13% | 15% |
| Q2 2025 | 34% | 12% | 14% |
| Q3 2025 | 37% | 12% | 16% |
| Q4 2025 | 40% | 14% | 18% |
| Q1 2026 | 43% | 16% | 21% |
| Q2 2026 | 46% | 18% | 24% |
We found that the open-weight share of production workloads fell from 15% in Q4 2024 to 12% in Q2 and Q3 2025, which matches the decline from 19% to 13% in open-source workload share that Menlo Ventures reported in its July 2025 market update. Respondents who pulled back in that period most often cited the gap between Llama 4 and the leading closed models.
Our data showed that the share of organizations running open-weight models in production rose in every quarter except Q2 2025, gaining 15 percentage points over the two-year window. Organizations kept existing open-weight deployments running even when they moved new workloads to closed APIs, so adoption held up better than workload share.
We observed the fastest growth in the most recent three quarters, when the open-weight token share climbed from 16% to 24% as newer Qwen, DeepSeek, and gpt-oss releases narrowed the quality gap with closed models. The open-weight share of workloads rose by 6 percentage points over the same period, from 12% to 18%.
US vs Chinese Open-Weight Model Share by Industry
In the table below, we break adoption down by industry and show two measures of model origin: the share of each industry's open-weight tokens that run on Chinese-origin models such as Qwen, DeepSeek, and Kimi, and the share of respondents with a formal policy restricting those models.
The US vs Chinese Open-Weight Model Share by Industry, 2026
| Industry | Respondents | Running Open-Weight Models in Production | Chinese-Origin Share of Open-Weight Tokens | Formally Restricting Chinese-Origin Models |
|---|---|---|---|---|
| Technology and software | 142 | 62% | 41% | 22% |
| Financial services | 118 | 44% | 12% | 58% |
| Healthcare and life sciences | 86 | 42% | 15% | 47% |
| Retail and ecommerce | 92 | 43% | 36% | 19% |
| Manufacturing | 104 | 42% | 29% | 33% |
| Professional services | 98 | 35% | 21% | 38% |
| All respondents | 640 | 46% | 28% | 36% |
Compliance policy shapes model origin more than any technical factor we measured. The four industries with the highest restriction rates, led by financial services at 58%, also have the four lowest Chinese-origin token shares. Technology companies show the opposite pattern: 62% run open-weight models in production and 41% of their open-weight tokens go to Chinese-origin models. Most restrictions we reviewed apply to hosted APIs operated from China and to models without a completed security review, and a minority ban Chinese-origin weights outright even when they run on the company's own hardware. Respondents in regulated industries often told us that legal review set the pace of adoption, ahead of model quality. Several financial services firms reported approved lists containing only US and European models, including Llama, Gemma, gpt-oss, and Mistral.
Enterprise usage still trails the developer market. OpenRouter data reported by Forkast in July 2026 put Chinese-origin models at 46% of tokens routed through the platform in mid-2026, and at no less than 30% in any week since February 8, 2026. A year earlier, the OpenRouter and Andreessen Horowitz token study found Chinese open-source models averaging about 13% of weekly platform tokens between November 2024 and November 2025. Menlo Ventures estimated in December 2025 that Chinese open-source models held roughly 10% of enterprise open-source usage, so our 28% figure points to a sharp rise inside companies during 2026. Our enterprise figure remains well below the 46% developer-platform figure, since restrictive policies in financial services and healthcare hold down the average.
Open-Weight Model Adoption by Use Case
In the table below, we show which tasks the 294 organizations running open-weight models in production assign to them, the open-weight share of tokens within each task, and the median cost savings respondents reported against the closed API they replaced or benchmarked. Respondents could select more than one use case.
The Open-Weight Model Adoption by Use Case, 2026
| Use Case | Share of Open-Weight Adopters | Open-Weight Share of Use-Case Tokens | Median Cost Savings vs Closed APIs |
|---|---|---|---|
| Document extraction and classification | 58% | 41% | 72% |
| Retrieval and internal search | 52% | 33% | 64% |
| Summarization and content generation | 47% | 30% | 67% |
| Customer support automation | 36% | 22% | 57% |
| Coding assistance | 29% | 11% | 38% |
| Agentic workflows | 21% | 9% | 31% |
We found that document extraction and classification is the leading open-weight use case, used by 58% of adopters, with 41% of its tokens on open-weight models and the highest median savings at 72%.
Our data showed that agentic workflows and coding assistance remain closed-model territory, with open-weight shares of 9% and 11%, as teams prioritize reliability on multi-step tool use over per-token price.
We observed that savings shrink as tasks become more complex: the three use cases with savings above 60% are all bounded text tasks, while the two with savings below 40% are coding and agents, where teams often need a larger model, longer outputs, or repeated attempts to match closed-model quality. Vercel reported in July 2026 that open-weight models on its platform cost roughly one-tenth of the average token price, so gross price gaps are far wider than the net savings companies report.
Open-Weight Model Deployment Methods and Cost Savings
In the table below, we group the 294 production adopters by their primary deployment method. Median cost savings compare total monthly inference cost, including hosting but excluding engineering staff, with the quoted price of an equivalent closed API.
The Open-Weight Model Deployment Methods and Cost Savings, 2026
| Primary Deployment Method | Adopters | Share of Adopters | Median Cost Savings vs Closed APIs | Chinese-Origin Share of Open-Weight Tokens |
|---|---|---|---|---|
| Self-hosted in a cloud VPC | 100 | 34% | 61% | 31% |
| Cloud marketplace | 85 | 29% | 52% | 19% |
| Self-hosted on premises | 56 | 19% | 48% | 22% |
| Third-party inference provider | 53 | 18% | 68% | 43% |
| All adopters | 294 | 100% | 57% | 28% |
Self-hosting in a cloud VPC is the most common setup, used by 34% of adopters, and cloud marketplaces such as Amazon Bedrock, Google Vertex AI, and Azure AI Foundry follow at 29%. Marketplaces deliver lower savings, at 52%, but they let companies buy open-weight inference under existing cloud contracts and security reviews, which helps explain their low 19% Chinese-origin share. Third-party inference providers such as Together AI, Fireworks, and Groq deliver the highest median savings at 68% and carry the highest Chinese-origin share at 43%, since they typically list new Qwen and DeepSeek releases within days. These providers also require the least infrastructure work, which makes them the usual starting point for companies with small platform teams.
On-premises deployments show the lowest savings at 48% because most organizations size their hardware for peak demand and leave it underused the rest of the day. The supply of models to deploy keeps widening: Hugging Face counted 151,448 derivative models built on Qwen by August 2026, against 82,506 built on Google's models.
Requesting a Copy of This Report
If you would like a PDF copy of this report or the underlying survey data, you can reach out here.
Sources
- Open-Weight Model Adoption Study, AI Industry Reviews, September 2026, New York, New York.
- 2025 Mid-Year LLM Market Update: Foundation Model Landscape + Economics, Menlo Ventures (Tim Tully, Joff Redfern, Deedy Das and Derek Xiao), July 2025, Menlo Park, California. https://menlovc.com/perspective/2025-mid-year-llm-market-update/
- 2025: The State of Generative AI in the Enterprise, Menlo Ventures (Tim Tully, Joff Redfern, Deedy Das and Derek Xiao), December 2025, Menlo Park, California. https://menlovc.com/perspective/2025-the-state-of-generative-ai-in-the-enterprise/
- Leaders, Gainers and Unexpected Winners in the Enterprise AI Arms Race, Andreessen Horowitz (Sarah Wang, Justin Kahl and Shangda Xu), January 2026, Menlo Park, California. https://a16z.com/leaders-gainers-and-unexpected-winners-in-the-enterprise-ai-arms-race/
- State of AI: An Empirical 100 Trillion Token Study with OpenRouter, OpenRouter and Andreessen Horowitz (Malika Aubakirova, Alex Atallah, Chris Clark, Justin Summerville and Anjney Midha), December 2025, New York, New York. https://openrouter.ai/assets/State-of-AI.pdf
- Chinese AI Models Now Capture Up to 46% of US Enterprise Token Usage, Yahoo Finance (Forkast News), July 2026, New York, New York. https://finance.yahoo.com/technology/ai/articles/chinese-ai-models-now-capture-020440715.html
- Chinese Open-Weight AI Models Gaining Ground as Enterprise Adoption Accelerates, Report, Computing (Dev Kundaliya), July 2026, London, United Kingdom. https://www.computing.co.uk/news/2026/ai/chinese-open-weight-ai-models-gaining-ground-as-enterprise-adoption-accelerates-report
- State of Open Models: Summer 2026 Observations, Hugging Face (Adina Yakefu and Irene Solaiman), August 2026, New York, New York. https://huggingface.co/blog/state-of-open-models-summer-2026
