Home / Research / AI Inference Cost Statistics
Market research · September 16, 2026
AI Inference Cost Statistics: September 2026
This report looks at published prices for running AI models as of September 2026. Those prices are for inference, the step where a trained model answers a prompt, generates text, or otherwise does useful work.
We priced 164 commercially available models on published input, output, cached, and batch rates. Where a provider splits rates by context length, we used the rate that applies below 200,000 tokens. The blended rate in the tables weights input at 75 percent and output at 25 percent.
AI Inference Cost by Model Tier
Tiers follow capability and positioning. Several providers now serve frontier-class quality from models they price as mid-tier.
The AI Inference Cost by Model Tier, 2026
| Model Tier | Median Input ($/MTok) | Median Output ($/MTok) | Blended Rate | Models Tracked |
|---|---|---|---|---|
| Frontier reasoning | $6.40 | $32.00 | $12.80 | 14 |
| Frontier standard | $3.75 | $18.50 | $7.44 | 19 |
| Mid-tier general | $1.15 | $6.20 | $2.41 | 31 |
| Budget | $0.28 | $1.40 | $0.56 | 42 |
| Open-weight hosted | $0.17 | $0.58 | $0.27 | 36 |
| Small and edge (under 4B) | $0.06 | $0.22 | $0.10 | 22 |
Insights
- The blended rate of the frontier reasoning tier is 128x above the small-model tier, wider than the 94x spread measured in the same exercise a year earlier.
- Output tokens price at a median of 5.0 times input tokens, a ratio that has held within half a point across every tier and every quarter tracked.
- Seventy-eight of the 164 models, or 48 percent, price below one dollar per million blended tokens. That band held only 22 models at the start of 2025.
Blended Inference Cost per Million Tokens, Q1 2023 to Q3 2026
Headline prices for the best available model have barely moved in three years, which is why the cost of inference is so often described as flat. The more useful series holds capability constant and asks what the cheapest model meeting a fixed quality bar charges. In the table below, we show all three series, followed by the same data plotted on a logarithmic scale.
The Blended Inference Cost per Million Tokens, 2026
| Quarter | Frontier Tier | Mid Tier | GPT-4-Class Capability |
|---|---|---|---|
| Q1 2023 | $37.50 | $8.60 | $37.50 |
| Q2 2023 | $34.20 | $7.40 | $30.20 |
| Q3 2023 | $31.40 | $6.20 | $18.40 |
| Q4 2023 | $28.60 | $5.10 | $9.60 |
| Q1 2024 | $24.80 | $4.30 | $5.40 |
| Q2 2024 | $21.30 | $3.60 | $3.10 |
| Q3 2024 | $18.60 | $3.10 | $1.60 |
| Q4 2024 | $16.20 | $2.85 | $0.97 |
| Q1 2025 | $14.70 | $2.60 | $0.61 |
| Q2 2025 | $12.40 | $2.45 | $0.44 |
| Q3 2025 | $10.90 | $2.52 | $0.33 |
| Q4 2025 | $11.60 | $2.38 | $0.26 |
| Q1 2026 | $13.20 | $2.44 | $0.19 |
| Q2 2026 | $12.60 | $2.35 | $0.17 |
| Q3 2026 | $12.80 | $2.41 | $0.14 |
The capability series fell by a factor of 268 over the period, an average of roughly 5x per year, while the frontier series fell by 71 percent through the third quarter of 2025 and then reversed. Both facts follow from the same commercial logic. Providers price the frontier at whatever the newest capability will bear, so that line tracks willingness to pay for new capability more closely than cost of production. Once a capability is two generations old, distilled and open-weight models compete on price alone, and the floor collapses. The 2026 uptick in the frontier line is the first sustained rise recorded in this series, and it coincides with reasoning models that bill for far more output tokens per answer than their predecessors did.
Inference Cost by Workload Type
A rate per million tokens tells a buyer very little without a token count attached to it. In the table below, we report the median token consumption measured for six common production workloads, along with what each one costs at mid-tier and budget-tier rates.
The Inference Cost by Workload Type, 2026
| Workload | Median Input Tokens | Median Output Tokens | Cost at Mid-Tier | Cost at Budget Tier |
|---|---|---|---|---|
| Single chat reply | 1,200 | 400 | $0.004 | $0.001 |
| Support answer with retrieval | 6,500 | 350 | $0.010 | $0.002 |
| Twenty-page document summary | 14,000 | 900 | $0.022 | $0.005 |
| Reasoning query with visible answer | 2,000 | 8,500 | $0.055 | $0.013 |
| Batch classification of 1,000 records | 420,000 | 30,000 | $0.670 | $0.160 |
| Fifty-turn agentic coding session | 1,050,000 | 42,000 | $1.470 | $0.350 |
The agentic session costs roughly 370 times what a single chat reply costs, and 96 percent of its tokens, carrying 82 percent of its cost, are input. The cause is structural: an agent resends its accumulated context on every turn, so a fifty-turn session pays for the same files and the same conversation history dozens of times over. In our sample, doubling the number of turns in a session raised its cost by a factor of 3.4, well above a linear doubling, and two runs of an identical task on an identical agent differed in cost by as much as 26 times, driven by how many times the agent chose to re-read its working set. Any organization forecasting inference spend from an average cost per request will underestimate agentic workloads badly.
Discounts by Cost Reduction Mechanism
Every major provider now sells the same tokens at several prices depending on how the request is submitted. In the table below, we quantify the median discount each mechanism delivered across the providers offering it.
The Discounts by Cost Reduction Mechanism, 2026
| Mechanism | Median Discount | Applies To | Providers Offering |
|---|---|---|---|
| Cached input reads | -88% | Repeated input tokens | 81% |
| Asynchronous batch submission | -50% | Input and output | 74% |
| Off-peak window pricing | -42% | Input and output | 23% |
| Committed throughput contract | -31% | Input and output | 56% |
| Routing to a lower tier | -71% | Blended workload cost | Not provider dependent |
Insights
- Cached input is the single largest lever available, and the one most often left unused: only 38 percent of the production deployments surveyed had structured their prompts so that the stable portion is at the front where it can be cached.
- Batch submission and cached reads are additive at most providers, producing an effective discount of 94 percent on repeated input tokens submitted asynchronously.
- Routing delivers the largest blended saving of any mechanism. The workloads best suited to it, including classification, extraction, and summarization, are also the highest-volume ones.
Self-Hosting Break-Even by API Tier
The alternative to paying per token is renting hardware and serving a model directly. In the table below, we solve for the monthly token volume at which an eight-GPU H100 node, rented at the median specialist cloud rate of $3.28 per GPU-hour and costing $18,900 per month, matches the cost of buying the same tokens from an API.
The Self-Hosting Break-Even by API Tier, 2026
| API Tier | Blended Rate | Break-Even Volume | Equivalent Chat Replies per Day |
|---|---|---|---|
| Frontier reasoning | $12.80 | 1.48 billion tokens | 30,800 |
| Frontier standard | $7.44 | 2.54 billion tokens | 52,900 |
| Mid-tier general | $2.41 | 7.84 billion tokens | 163,000 |
| Budget | $0.56 | 33.8 billion tokens | 704,000 |
| Open-weight hosted | $0.27 | 70.0 billion tokens | 1,458,000 |
The break-even volumes above count only the rented hardware. They exclude the engineering time to deploy, quantize, monitor, and update a served model, which in our sample added between 30 and 100 percent to the true cost for teams under twenty engineers. They also assume a node kept near full utilization, which almost never describes real traffic: a workload with a weekday peak and a quiet weekend runs at 30 to 40 percent average utilization unless it is deliberately batched. Self-hosting pays off against frontier pricing at volumes a mid-sized company can reach. It rarely pays off against open-weight hosted pricing, because the providers of those endpoints already run the same hardware at utilization rates a single tenant cannot match.
Requesting a Copy of This Report
If you would like a PDF copy of this report, or the full pricing dataset behind it, you can reach out here.
Sources
- AI Inference Pricing Study. AI Industry Reviews. September 2026. New York, New York.
- LLM Inference Prices Have Fallen Rapidly but Unequally Across Tasks, Epoch AI, 2026, Oakland, California. https://epoch.ai/data-insights/llm-inference-price-trends
- Welcome to LLMflation, Andreessen Horowitz, November 2024, Menlo Park, California. https://a16z.com/llmflation-llm-inference-cost/
- LLM API Pricing 2026: OpenAI, Gemini, Claude and Grok, IntuitionLabs, August 2026, San Francisco, California. https://intuitionlabs.ai/articles/llm-api-pricing-comparison-2025
- AI Inference Cost Statistics 2026, Axis Intelligence, July 2026, Austin, Texas. https://axis-intelligence.com/ai-inference-cost-statistics/
- The Hidden Cost Driver in Agentic Coding Sessions, Vantage, April 2026, New York, New York. https://www.vantage.sh/blog/agentic-coding-costs
