Home / Research / Open-Weight Model Adoption Statistics

Market research · September 24, 2026

Open-Weight Model Adoption Statistics: 2026 Report

This report looks at how companies use open-weight models such as Meta's Llama, Alibaba's Qwen, DeepSeek, Mistral, Google's Gemma, and OpenAI's gpt-oss alongside closed APIs. Open-weight models are models whose weights are publicly released, so organizations can download them and run inference on their own hardware or through a hosting provider.

We surveyed 640 organizations with 100 or more employees in North America and Europe between March and August 2026, collecting quarterly usage figures from Q3 2024 to Q2 2026. We combined our data with published figures from Menlo Ventures, OpenRouter, Andreessen Horowitz, Hugging Face, and Vercel.

We found that 46% of organizations ran at least one open-weight model in production at the end of Q2 2026, up from 31% in Q3 2024. Open-weight models handled 24% of those organizations' LLM tokens but only 18% of their production workloads, because they are concentrated in high-volume tasks such as document extraction. Chinese-origin models accounted for 28% of open-weight tokens, while 36% of organizations formally restrict them.

Open-Weight Model Adoption by Company Size

In the table below, we show the share of organizations running open-weight models in production, along with the open-weight share of production workloads and of monthly tokens, as of the end of Q2 2026. We count an organization as a production user when at least one open-weight model serves a live customer-facing or employee-facing workload.

The Open-Weight Model Adoption by Company Size, 2026

Company SizeRespondentsRunning Open-Weight Models in ProductionOpen-Weight Share of Production WorkloadsOpen-Weight Share of Tokens
100 to 999 employees22838%21%27%
1,000 to 9,999 employees24647%17%24%
10,000+ employees16656%14%21%
All respondents64046%18%24%
Grouped bar chart of open-weight production adoption, workload share, and token share by company size in 2026.
The Open-Weight Model Adoption by Company Size, September 2026

Adoption rises with company size, while depth of use falls. Organizations with 10,000 or more employees were 18 percentage points more likely than those with 100 to 999 employees to run an open-weight model in production, yet open-weight models handled 14% of their workloads against 21% at the smaller firms. Large enterprises tend to add an open-weight model for one or two high-volume tasks while keeping closed APIs as the default for everything else. Smaller companies that adopt open-weight models more often build a larger share of their stack on them, frequently through a single engineering team that standardizes on one model family. The token column tells the same story from another angle. At every company size, the open-weight share of tokens exceeds the open-weight share of workloads, by 6 to 7 percentage points, which means the workloads moved to open-weight models are heavier than average. Respondents described these as batch jobs over large document sets, where per-token price dominates the purchasing decision.

Our 24% token share is well above the 11% open-source share of enterprise LLM usage that Menlo Ventures reported in December 2025. Our figure is measured half a year later and includes mid-sized companies, which Menlo's enterprise sample weights less heavily. Closed providers still dominate: the Andreessen Horowitz CIO survey of 100 Global 2000 companies, published in January 2026, found OpenAI models in production at 78% of respondents. Open-weight models have become a common second supplier inside large companies, while the primary relationship for most enterprises remains a closed API provider.

Open-Weight Model Adoption Trend, Q3 2024 to Q2 2026

The table below tracks the same three measures for our full panel at the end of each quarter over the past two years.

The Open-Weight Model Adoption Trend, Q3 2024 to Q2 2026

QuarterRunning Open-Weight Models in ProductionOpen-Weight Share of Production WorkloadsOpen-Weight Share of Tokens
Q3 202431%14%13%
Q4 202433%15%14%
Q1 202535%13%15%
Q2 202534%12%14%
Q3 202537%12%16%
Q4 202540%14%18%
Q1 202643%16%21%
Q2 202646%18%24%
Line chart of open-weight production adoption, workload share, and token share from Q3 2024 to Q2 2026.
The Open-Weight Model Adoption Trend, Q3 2024 to Q2 2026

We found that the open-weight share of production workloads fell from 15% in Q4 2024 to 12% in Q2 and Q3 2025, which matches the decline from 19% to 13% in open-source workload share that Menlo Ventures reported in its July 2025 market update. Respondents who pulled back in that period most often cited the gap between Llama 4 and the leading closed models.

Our data showed that the share of organizations running open-weight models in production rose in every quarter except Q2 2025, gaining 15 percentage points over the two-year window. Organizations kept existing open-weight deployments running even when they moved new workloads to closed APIs, so adoption held up better than workload share.

We observed the fastest growth in the most recent three quarters, when the open-weight token share climbed from 16% to 24% as newer Qwen, DeepSeek, and gpt-oss releases narrowed the quality gap with closed models. The open-weight share of workloads rose by 6 percentage points over the same period, from 12% to 18%.

US vs Chinese Open-Weight Model Share by Industry

In the table below, we break adoption down by industry and show two measures of model origin: the share of each industry's open-weight tokens that run on Chinese-origin models such as Qwen, DeepSeek, and Kimi, and the share of respondents with a formal policy restricting those models.

The US vs Chinese Open-Weight Model Share by Industry, 2026

IndustryRespondentsRunning Open-Weight Models in ProductionChinese-Origin Share of Open-Weight TokensFormally Restricting Chinese-Origin Models
Technology and software14262%41%22%
Financial services11844%12%58%
Healthcare and life sciences8642%15%47%
Retail and ecommerce9243%36%19%
Manufacturing10442%29%33%
Professional services9835%21%38%
All respondents64046%28%36%
Grouped bar chart of open-weight production adoption, Chinese-origin token share, and restriction rates by industry in 2026.
The US vs Chinese Open-Weight Model Share by Industry, September 2026

Compliance policy shapes model origin more than any technical factor we measured. The four industries with the highest restriction rates, led by financial services at 58%, also have the four lowest Chinese-origin token shares. Technology companies show the opposite pattern: 62% run open-weight models in production and 41% of their open-weight tokens go to Chinese-origin models. Most restrictions we reviewed apply to hosted APIs operated from China and to models without a completed security review, and a minority ban Chinese-origin weights outright even when they run on the company's own hardware. Respondents in regulated industries often told us that legal review set the pace of adoption, ahead of model quality. Several financial services firms reported approved lists containing only US and European models, including Llama, Gemma, gpt-oss, and Mistral.

Enterprise usage still trails the developer market. OpenRouter data reported by Forkast in July 2026 put Chinese-origin models at 46% of tokens routed through the platform in mid-2026, and at no less than 30% in any week since February 8, 2026. A year earlier, the OpenRouter and Andreessen Horowitz token study found Chinese open-source models averaging about 13% of weekly platform tokens between November 2024 and November 2025. Menlo Ventures estimated in December 2025 that Chinese open-source models held roughly 10% of enterprise open-source usage, so our 28% figure points to a sharp rise inside companies during 2026. Our enterprise figure remains well below the 46% developer-platform figure, since restrictive policies in financial services and healthcare hold down the average.

Open-Weight Model Adoption by Use Case

In the table below, we show which tasks the 294 organizations running open-weight models in production assign to them, the open-weight share of tokens within each task, and the median cost savings respondents reported against the closed API they replaced or benchmarked. Respondents could select more than one use case.

The Open-Weight Model Adoption by Use Case, 2026

Use CaseShare of Open-Weight AdoptersOpen-Weight Share of Use-Case TokensMedian Cost Savings vs Closed APIs
Document extraction and classification58%41%72%
Retrieval and internal search52%33%64%
Summarization and content generation47%30%67%
Customer support automation36%22%57%
Coding assistance29%11%38%
Agentic workflows21%9%31%
Horizontal grouped bar chart of open-weight adopter share and use-case token share by use case in 2026.
The Open-Weight Model Adoption by Use Case, September 2026

We found that document extraction and classification is the leading open-weight use case, used by 58% of adopters, with 41% of its tokens on open-weight models and the highest median savings at 72%.

Our data showed that agentic workflows and coding assistance remain closed-model territory, with open-weight shares of 9% and 11%, as teams prioritize reliability on multi-step tool use over per-token price.

We observed that savings shrink as tasks become more complex: the three use cases with savings above 60% are all bounded text tasks, while the two with savings below 40% are coding and agents, where teams often need a larger model, longer outputs, or repeated attempts to match closed-model quality. Vercel reported in July 2026 that open-weight models on its platform cost roughly one-tenth of the average token price, so gross price gaps are far wider than the net savings companies report.

Open-Weight Model Deployment Methods and Cost Savings

In the table below, we group the 294 production adopters by their primary deployment method. Median cost savings compare total monthly inference cost, including hosting but excluding engineering staff, with the quoted price of an equivalent closed API.

The Open-Weight Model Deployment Methods and Cost Savings, 2026

Primary Deployment MethodAdoptersShare of AdoptersMedian Cost Savings vs Closed APIsChinese-Origin Share of Open-Weight Tokens
Self-hosted in a cloud VPC10034%61%31%
Cloud marketplace8529%52%19%
Self-hosted on premises5619%48%22%
Third-party inference provider5318%68%43%
All adopters294100%57%28%
Grouped bar chart of adopter share and median cost savings by open-weight deployment method in 2026.
The Open-Weight Model Deployment Methods and Cost Savings, September 2026

Self-hosting in a cloud VPC is the most common setup, used by 34% of adopters, and cloud marketplaces such as Amazon Bedrock, Google Vertex AI, and Azure AI Foundry follow at 29%. Marketplaces deliver lower savings, at 52%, but they let companies buy open-weight inference under existing cloud contracts and security reviews, which helps explain their low 19% Chinese-origin share. Third-party inference providers such as Together AI, Fireworks, and Groq deliver the highest median savings at 68% and carry the highest Chinese-origin share at 43%, since they typically list new Qwen and DeepSeek releases within days. These providers also require the least infrastructure work, which makes them the usual starting point for companies with small platform teams.

On-premises deployments show the lowest savings at 48% because most organizations size their hardware for peak demand and leave it underused the rest of the day. The supply of models to deploy keeps widening: Hugging Face counted 151,448 derivative models built on Qwen by August 2026, against 82,506 built on Google's models.

Requesting a Copy of This Report

If you would like a PDF copy of this report or the underlying survey data, you can reach out here.

Sources