Home / Research / AI Inference Cost Statistics

Market research · September 16, 2026

AI Inference Cost Statistics: September 2026

This report looks at published prices for running AI models as of September 2026. Those prices are for inference, the step where a trained model answers a prompt, generates text, or otherwise does useful work.

We priced 164 commercially available models on published input, output, cached, and batch rates. Where a provider splits rates by context length, we used the rate that applies below 200,000 tokens. The blended rate in the tables weights input at 75 percent and output at 25 percent.

AI Inference Cost by Model Tier

Tiers follow capability and positioning. Several providers now serve frontier-class quality from models they price as mid-tier.

The AI Inference Cost by Model Tier, 2026

Model TierMedian Input ($/MTok)Median Output ($/MTok)Blended RateModels Tracked
Frontier reasoning$6.40$32.00$12.8014
Frontier standard$3.75$18.50$7.4419
Mid-tier general$1.15$6.20$2.4131
Budget$0.28$1.40$0.5642
Open-weight hosted$0.17$0.58$0.2736
Small and edge (under 4B)$0.06$0.22$0.1022

Insights

Blended Inference Cost per Million Tokens, Q1 2023 to Q3 2026

Headline prices for the best available model have barely moved in three years, which is why the cost of inference is so often described as flat. The more useful series holds capability constant and asks what the cheapest model meeting a fixed quality bar charges. In the table below, we show all three series, followed by the same data plotted on a logarithmic scale.

The Blended Inference Cost per Million Tokens, 2026

QuarterFrontier TierMid TierGPT-4-Class Capability
Q1 2023$37.50$8.60$37.50
Q2 2023$34.20$7.40$30.20
Q3 2023$31.40$6.20$18.40
Q4 2023$28.60$5.10$9.60
Q1 2024$24.80$4.30$5.40
Q2 2024$21.30$3.60$3.10
Q3 2024$18.60$3.10$1.60
Q4 2024$16.20$2.85$0.97
Q1 2025$14.70$2.60$0.61
Q2 2025$12.40$2.45$0.44
Q3 2025$10.90$2.52$0.33
Q4 2025$11.60$2.38$0.26
Q1 2026$13.20$2.44$0.19
Q2 2026$12.60$2.35$0.17
Q3 2026$12.80$2.41$0.14
Log-scale line chart of blended inference cost per million tokens from Q1 2023 to Q3 2026 for frontier tier, mid tier, and GPT-4-class capability.
The Blended Inference Cost per Million Tokens, September 2026 (logarithmic scale)

The capability series fell by a factor of 268 over the period, an average of roughly 5x per year, while the frontier series fell by 71 percent through the third quarter of 2025 and then reversed. Both facts follow from the same commercial logic. Providers price the frontier at whatever the newest capability will bear, so that line tracks willingness to pay for new capability more closely than cost of production. Once a capability is two generations old, distilled and open-weight models compete on price alone, and the floor collapses. The 2026 uptick in the frontier line is the first sustained rise recorded in this series, and it coincides with reasoning models that bill for far more output tokens per answer than their predecessors did.

Inference Cost by Workload Type

A rate per million tokens tells a buyer very little without a token count attached to it. In the table below, we report the median token consumption measured for six common production workloads, along with what each one costs at mid-tier and budget-tier rates.

The Inference Cost by Workload Type, 2026

WorkloadMedian Input TokensMedian Output TokensCost at Mid-TierCost at Budget Tier
Single chat reply1,200400$0.004$0.001
Support answer with retrieval6,500350$0.010$0.002
Twenty-page document summary14,000900$0.022$0.005
Reasoning query with visible answer2,0008,500$0.055$0.013
Batch classification of 1,000 records420,00030,000$0.670$0.160
Fifty-turn agentic coding session1,050,00042,000$1.470$0.350

The agentic session costs roughly 370 times what a single chat reply costs, and 96 percent of its tokens, carrying 82 percent of its cost, are input. The cause is structural: an agent resends its accumulated context on every turn, so a fifty-turn session pays for the same files and the same conversation history dozens of times over. In our sample, doubling the number of turns in a session raised its cost by a factor of 3.4, well above a linear doubling, and two runs of an identical task on an identical agent differed in cost by as much as 26 times, driven by how many times the agent chose to re-read its working set. Any organization forecasting inference spend from an average cost per request will underestimate agentic workloads badly.

Discounts by Cost Reduction Mechanism

Every major provider now sells the same tokens at several prices depending on how the request is submitted. In the table below, we quantify the median discount each mechanism delivered across the providers offering it.

The Discounts by Cost Reduction Mechanism, 2026

MechanismMedian DiscountApplies ToProviders Offering
Cached input reads-88%Repeated input tokens81%
Asynchronous batch submission-50%Input and output74%
Off-peak window pricing-42%Input and output23%
Committed throughput contract-31%Input and output56%
Routing to a lower tier-71%Blended workload costNot provider dependent

Insights

Self-Hosting Break-Even by API Tier

The alternative to paying per token is renting hardware and serving a model directly. In the table below, we solve for the monthly token volume at which an eight-GPU H100 node, rented at the median specialist cloud rate of $3.28 per GPU-hour and costing $18,900 per month, matches the cost of buying the same tokens from an API.

The Self-Hosting Break-Even by API Tier, 2026

API TierBlended RateBreak-Even VolumeEquivalent Chat Replies per Day
Frontier reasoning$12.801.48 billion tokens30,800
Frontier standard$7.442.54 billion tokens52,900
Mid-tier general$2.417.84 billion tokens163,000
Budget$0.5633.8 billion tokens704,000
Open-weight hosted$0.2770.0 billion tokens1,458,000

The break-even volumes above count only the rented hardware. They exclude the engineering time to deploy, quantize, monitor, and update a served model, which in our sample added between 30 and 100 percent to the true cost for teams under twenty engineers. They also assume a node kept near full utilization, which almost never describes real traffic: a workload with a weekday peak and a quiet weekend runs at 30 to 40 percent average utilization unless it is deliberately batched. Self-hosting pays off against frontier pricing at volumes a mid-sized company can reach. It rarely pays off against open-weight hosted pricing, because the providers of those endpoints already run the same hardware at utilization rates a single tenant cannot match.

Requesting a Copy of This Report

If you would like a PDF copy of this report, or the full pricing dataset behind it, you can reach out here.

Sources