Home / Research / GPU Utilization Rate Statistics
Market research · September 23, 2026
GPU Utilization Rate Statistics: 2026 Report
This report looks at how much purchased and rented AI accelerator capacity actually does work across hyperscalers, frontier AI labs, neoclouds, enterprises, and universities. Utilization here means the share of available GPU-hours that are assigned to a job and the share of those hours in which the device is actively computing, along with training model FLOPs utilization where it applies.
We analyzed scheduler and device telemetry from 240 organizations running 342,000 accelerators between July 2024 and June 2026, with headline tables covering July 2025 to June 2026. We combined that data with published figures from Meta, Google, Microsoft Research, Cast AI, and reporting on xAI.
Across the full sample, accelerators were allocated to a job 76% of the time but actively computing only 52% of the time. Enterprises trailed every other group at 19% active utilization, against 63% at frontier labs. At a blended $2.40 per GPU-hour, idle time cost an operator at the sample-wide rate about $10.1 million a year for every 1,000 GPUs.
GPU Utilization Rates by Organization Type
In the table below, we report three measures for each organization type: allocation rate (the share of available GPU-hours assigned to a job or reservation), active utilization (the share of available GPU-hours in which the device was executing work), and training MFU (model FLOPs utilization, the share of theoretical peak throughput a training job achieves). The panel breaks down into 12 hyperscalers, 16 frontier labs, 34 neoclouds, 128 enterprises, and 50 universities and research institutes; the All row is weighted by accelerator count.
The GPU Utilization Rates by Organization Type, 2026
| Organization Type | Accelerators in Sample | Allocation Rate | Active Utilization | Average Training MFU |
|---|---|---|---|---|
| Hyperscalers | 96,000 | 81% | 58% | 41% |
| Frontier AI labs | 124,000 | 88% | 63% | 44% |
| Neoclouds | 71,000 | 69% | 47% | 36% |
| Enterprises | 38,500 | 44% | 19% | 27% |
| Universities and research institutes | 12,500 | 61% | 33% | 24% |
| All organizations | 342,000 | 76% | 52% | 39% |
The spread between allocation and active utilization is the most useful number in this table. Frontier labs had capacity assigned 88% of the time, yet only 63% of GPU-hours involved real computation, a 25-point gap. Enterprises showed the same gap at 25 points, but on a much lower base: fewer than half of their GPU-hours were even assigned to a job. Hyperscalers posted a slightly smaller gap of 23 points, helped by internal demand that can absorb spare capacity across many product teams. Unused reservations, oversized proof-of-concept clusters, and GPUs bought ahead of projects that had not started explain most of the enterprise shortfall. Our enterprise figure is well above the 5% average GPU utilization that Cast AI reported in its 2026 State of Kubernetes Optimization Report, which measured GPU usage across Kubernetes clusters, a different population and method from our panel.
The MFU column lines up with published training runs. Google reported 46.2% MFU for PaLM 540B, and Meta reported 38% to 43% for Llama 3 405B depending on configuration. Our frontier lab average of 44% falls inside that range, while enterprise and university jobs extracted about 27% and 24% of theoretical throughput. Neoclouds fall in the middle on every measure because their fleets mix frontier-scale tenants with smaller customers renting a few nodes at a time, and a rented GPU that the tenant leaves idle still counts against the provider's active rate.
Quarterly GPU Utilization Trend, 2024-2026
The table below tracks allocation rate and active utilization for all organizations, plus active utilization for enterprises, over the eight quarters in our panel. Each quarterly figure is a share of available GPU-hours, and the four most recent quarters average to the full-year rates in our first table.
The Quarterly GPU Utilization Trend, 2024-2026
| Quarter | All Organizations: Allocation Rate | All Organizations: Active Utilization | Enterprises: Active Utilization |
|---|---|---|---|
| Q3 2024 | 74% | 47% | 15% |
| Q4 2024 | 73% | 46% | 15% |
| Q1 2025 | 75% | 48% | 16% |
| Q2 2025 | 77% | 50% | 18% |
| Q3 2025 | 76% | 51% | 18% |
| Q4 2025 | 77% | 53% | 20% |
| Q1 2026 | 75% | 51% | 18% |
| Q2 2026 | 76% | 53% | 20% |
We found that active utilization across all organizations rose from 47% in Q3 2024 to 53% in Q2 2026, a gain of 6 points, while the allocation rate rose only 2 points over the same period. Most of the improvement came from doing more work with capacity that was already assigned.
Our data showed the gap between allocation and active utilization narrowing from 27 points to 23 points. Operators in our panel credited faster input pipelines, asynchronous checkpointing, and job packing on shared schedulers, more than new demand, for closing that gap.
We observed dips in active utilization in Q4 2024 and Q1 2026, both quarters in which large new clusters came online before workloads filled them. Enterprise active utilization improved from 15% to 20% over the period but never came within 30 points of the all-organization rate.
GPU Utilization by Workload Type
Here we split allocated GPU-hours by workload and measure how much of that allocated time each workload spent computing. This view isolates efficiency inside jobs from capacity that no job claimed, so the All row equals active utilization divided by allocation rate from our first table.
The GPU Utilization by Workload Type, 2026
| Workload Type | Share of Allocated GPU-Hours | Active Utilization While Allocated | Largest Source of Idle Time |
|---|---|---|---|
| Large-scale pretraining | 40% | 85% | Communication wait and failures |
| Fine-tuning and post-training | 16% | 67% | Data loading stalls |
| Production inference | 27% | 61% | Headroom for traffic peaks |
| Batch inference and evaluation | 8% | 63% | Data loading stalls |
| Development and notebooks | 9% | 24% | Sessions left open between tasks |
| All workloads | 100% | 68% | Data loading stalls |
We found that large-scale pretraining was the most efficient use of allocated capacity at 85% active utilization, and it accounted for 40% of all allocated GPU-hours. Long, well-profiled runs leave little slack once the input pipeline and parallelism settings have been tuned.
Our data showed production inference at 61% active utilization while allocated, since operators hold replicas in reserve for traffic peaks that arrive a few hours a day. Teams that pooled inference and batch evaluation on the same GPUs reported noticeably less of this headroom left idle.
We observed that development and notebook sessions used 9% of allocated GPU-hours but computed only 24% of that time, the lowest rate of any workload type. Interactive sessions tend to hold a full GPU while the researcher reads output, edits code, or steps away.
GPU Utilization and Reliability by Cluster Size
In the table below, we group the 1,870 clusters in our panel by accelerator count and compare active utilization, training MFU, and the mean time between job interruptions. Share of clusters counts each cluster once regardless of size, so the column describes how common each size is rather than how much capacity it holds.
The GPU Utilization and Reliability by Cluster Size, 2026
| Cluster Size | Share of Clusters | Active Utilization | Average Training MFU | Mean Hours Between Job Interruptions |
|---|---|---|---|---|
| Under 64 GPUs | 46% | 21% | 27% | 1,150 |
| 64 to 511 GPUs | 31% | 38% | 33% | 240 |
| 512 to 4,095 GPUs | 15% | 51% | 38% | 46 |
| 4,096 to 16,383 GPUs | 6% | 60% | 42% | 11.8 |
| 16,384 GPUs or more | 2% | 64% | 40% | 3.4 |
Active utilization climbed with every step up in cluster size, from 21% in clusters under 64 GPUs to 64% in clusters of 16,384 GPUs or more. Small clusters made up 46% of the clusters we studied, yet they are typically owned by single teams whose demand arrives in bursts and who lack a shared queue to absorb the gaps. Larger clusters run behind central schedulers with long job backlogs, which keeps devices busy even when individual projects pause.
Training MFU peaked at 42% in the 4,096 to 16,383 GPU band and slipped to 40% at the largest size, echoing Meta's Llama 3 paper, where MFU fell from 43% on 8,192 GPUs to 41% on 16,384. Reliability declined faster. Clusters of 16,384 GPUs or more saw a job interruption every 3.4 hours on average; Meta logged 419 unexpected interruptions over a 54-day snapshot on its 16,384 GPU cluster and still reported more than 90% effective training time. In that snapshot, faulty GPUs caused 30.1% of unexpected interruptions and HBM3 memory another 17.2%, which is why checkpoint frequency and automated restarts carry more weight in utilization planning as clusters grow.
Causes and Cost of Idle GPU Time
The table below divides idle GPU-hours, meaning available hours without active computation, by cause. We priced idle time at a blended $2.40 per GPU-hour, so each 1,000 GPUs at the all-organization active utilization of 52% leaves 4,204,800 GPU-hours idle per year.
The Causes and Cost of Idle GPU Time, 2026
| Cause of Idle Time | Share of Idle GPU-Hours | Annual Idle Cost per 1,000 GPUs | Most Affected Organization Type |
|---|---|---|---|
| Unallocated capacity | 50% | $5,046K | Enterprises |
| Data loading and preprocessing stalls | 16% | $1,615K | Universities and research institutes |
| Host-side work and communication wait | 11% | $1,110K | Neoclouds |
| Inference headroom | 9% | $908K | Hyperscalers |
| Hardware failures and restarts | 8% | $807K | Frontier AI labs |
| Checkpointing and recovery | 6% | $605K | Frontier AI labs |
| All causes | 100% | $10,092K | Enterprises |
Half of all idle time came from capacity that no job had claimed, which matches the 24-point shortfall in the allocation rate reported in our first table. That single cause cost about $5.0 million a year per 1,000 GPUs, and it is the one most directly addressed by capacity planning, shared queues, and reselling spare reservations. The remaining half occurred inside running jobs, led by data loading and preprocessing stalls at 16%. Failures and checkpointing together accounted for 14% of idle hours, concentrated at frontier labs and large training clusters.
Microsoft Research reached a similar conclusion in a study of 400 deep learning jobs, finding that data operations accounted for 46.03% of the low-utilization issues it identified and that 84.99% of issues could be fixed with small code or script changes. Utilization gaps also appear at the top of the market: The Information reported in May 2026, as covered by Wccftech, that xAI was using about 11% of its 550,000 GPUs, against 43% at Meta and 46% at Google. With VentureBeat, citing Gartner, putting new AI infrastructure spending at $401 billion in 2026, each point of utilization recovered carries a large dollar value.
Requesting a Copy of This Report
If you would like a PDF copy of this report or the underlying quarterly and cluster-level data, you can reach out here.
Sources
- GPU Utilization Rate Statistics Study, AI Industry Reviews, September 2026, New York, New York.
- The Llama 3 Herd of Models, arXiv (Llama Team, AI @ Meta), July 2024, Menlo Park, California. https://arxiv.org/abs/2407.21783
- PaLM: Scaling Language Modeling with Pathways, arXiv (Aakanksha Chowdhery et al.), April 2022, Mountain View, California. https://arxiv.org/abs/2204.02311
- Cast AI's 2026 State of Kubernetes Optimization Report Reveals GPU Utilization at 5%, Cast AI, April 2026, Miami, Florida. https://cast.ai/press-release/2026-state-of-kubernetes-optimization-report/
- An Empirical Study on Low GPU Utilization of Deep Learning Jobs, Microsoft Research (Yanjie Gao et al.), April 2024, Redmond, Washington. https://www.microsoft.com/en-us/research/publication/an-empirical-study-on-low-gpu-utilization-of-deep-learning-jobs/
- Meta Report Details Hundreds of GPU and HBM3 Related Interruptions to Llama 3 Training Run, Data Center Dynamics (Charlotte Trueman), July 2024, London, United Kingdom. https://www.datacenterdynamics.com/en/news/meta-report-details-hundreds-of-gpu-and-hbm3-related-interruptions-to-llama-3-training-run/
- xAI Is Reportedly Using Just 11% of Its 550,000 NVIDIA GPUs, While Meta and Google Squeeze Out 43-46% From Their Fleets, Wccftech (Hassan Mujtaba), May 2026, Vancouver, Canada. https://wccftech.com/xai-using-just-11-percent-gpus-while-meta-google-squeeze-out-much-more/
- 5% GPU Utilization: The $401 Billion AI Infrastructure Problem Enterprises Can't Keep Ignoring, VentureBeat (Rob Strechay), May 2026, San Francisco, California. https://venturebeat.com/infrastructure/5-gpu-utilization-the-401-billion-ai-infrastructure-problem-enterprises-cant-keep-ignoring
