Home / Research / AI Token Pricing Comparison

Market research · September 2026

AI Token Pricing Comparison: September 2026

Published rate cards miss token burn. Across the four models in this study, burn varies by more than 3x. Two models at the same $50.00 per million output differ by 22 percent on a long document and by 63 percent across a coding session.

The study covers GPT-6 Astra (OpenAI), Claude Fable 5.1 (Anthropic), Gemini 3.1 Pro (Google; API id gemini-3.1-pro-preview, shortened in this report), and Grok 4.6 (xAI), each at default reasoning through the vendor's first-party API. The 240-task suite is split evenly across agentic coding, long-document analysis, multi-step research, and structured extraction. Each call logged input, cached input, visible output, and billed reasoning, and was priced on published rates as of September 1, 2026. Rates are standard synchronous unless a table says otherwise.

What the study covers

Layer What it measures Measurement scope
Published price Price per million tokens on the published rate card for input, cached input, output, and long-context First-party API rate cards as of September 1, 2026; standard synchronous unless a table says otherwise
Token burn Visible output, billed reasoning, and total billed output per task 240-task suite split evenly across agentic coding, long-document analysis, multi-step research, and structured extraction
Session cost Cumulative spend across a multi-turn agentic coding run 50-turn session of 1,030,500 input tokens with 72 percent prefix cache hit from the third turn
Workload cost Cost of 1,000 executions of a defined production job Four workloads at 1,000 runs
Monthly spend Monthly API spend per developer, including the share from re-sent context and from system prompts and tools 34 engineering teams on one primary model

Published price is the rate. Token burn is the quantity billed to finish a task. Session cost, workload cost, and monthly spend apply both figures to larger units of work.

Published token prices

The table lists each model's published rate per million tokens, its context window, and its long-context rate. Google and xAI set the long-context threshold at 200,000 tokens and bill the higher rate against every token in the request once a prompt reaches that line. OpenAI publishes long-context rates without a published threshold; this study applied the same 200,000-token line.

Model Vendor Input Cached input Output Context window Long-context rate
GPT-6 Astra OpenAI $10.00 $1.00 $50.00 1.05M 2x input, 1.5x output
Claude Fable 5.1 Anthropic $10.00 $0.25 $50.00 1M None, flat to 1M
Gemini 3.1 Pro Google $2.00 $0.20 $12.00 1M $4.00 in, $18.00 out
Grok 4.6 xAI $2.00 $0.50 $6.00 500K $4.00 in, $12.00 out

Insights

Token consumption per task

The rate card sets the price of a token. The task sets how many tokens are billed. The table reports average output per task across the 240-task suite, split into visible output and billed reasoning, with a burn multiple indexed to GPT-6 Astra at 1.00x.

Model Visible output tokens Reasoning output tokens Total billed output Reasoning share Burn multiple
GPT-6 Astra 2,500 5,900 8,400 70.2% 1.00x
Gemini 3.1 Pro 3,100 6,600 9,700 68.0% 1.15x
Grok 4.6 3,400 9,700 13,100 74.0% 1.56x
Claude Fable 5.1 5,100 21,300 26,400 80.7% 3.14x
Stacked bar chart of billed output tokens per task, split into visible output and billed reasoning, for GPT-6 Astra, Gemini 3.1 Pro, Grok 4.6, and Claude Fable 5.1.
Billed output tokens per task, visible versus reasoning, 2026.

Insights

Cumulative session cost across a 50-turn agentic run

Agentic runs resend accumulated working context as input on each turn. The table tracks cumulative spend across a 50-turn coding session of 1,030,500 input tokens, with a 72 percent prefix cache hit from the third turn. OpenAI and Anthropic bill cache writes above base input, and those writes are priced in. Google and xAI publish no separate cache-write line.

Turn GPT-6 Astra Claude Fable 5.1 Gemini 3.1 Pro Grok 4.6
1 $0.09 $0.18 $0.02 $0.02
5 $0.40 $0.82 $0.09 $0.08
10 $0.78 $1.60 $0.18 $0.16
15 $1.23 $2.45 $0.28 $0.26
20 $1.72 $3.33 $0.40 $0.37
25 $2.28 $4.27 $0.52 $0.50
30 $2.91 $5.28 $0.65 $0.64
35 $3.58 $6.31 $0.80 $0.79
40 $4.32 $7.43 $0.95 $0.96
45 $5.12 $8.57 $1.12 $1.14
50 $5.97 $9.75 $1.30 $1.34
Line chart of cumulative session cost across a 50-turn agentic run for GPT-6 Astra, Claude Fable 5.1, Gemini 3.1 Pro, and Grok 4.6.
Cumulative session cost across a 50-turn agentic run, 2026.

The chart plots the same figures. The four session totals separate on the first turn and remain apart through turn 50. By turn 10 the spread is $1.44. By turn 50 it is $8.45, on a session that produced the same work in every case. Input grows from about 5,400 tokens on turn 1 to about 35,800 on turn 50. Caching reduces the input side of the bill and leaves per-turn output burn as the larger remaining variable. Gemini 3.1 Pro and Grok 4.6 stay within four cents for the run. A caching-disabled control on the same session raised session totals 71 to 106 percent.

Effective cost per 1,000 runs by workload

Rate and burn combine differently by the shape of the job. The table prices four workloads at 1,000 runs each, with the task held constant across models. The customer support reply and structured classification rows ran with reasoning disabled, so output length matches and only the rate card varies. Structured classification assumes batch wherever the vendor offers it.

Workload Input tokens per run GPT-6 Astra Claude Fable 5.1 Gemini 3.1 Pro Grok 4.6
Customer support reply 2,400 $24.72 $23.73 $5.30 $4.26
Long document analysis 120,000 $1,023.60 $1,248.60 $207.79 $208.01
Full repository review 340,000 $4,540.50 $2,851.25 $922.69 $1,003.18
Structured classification 1,800 $8.00 $7.46 $1.76 $2.70
Grouped bar chart of relative cost per workload, indexed to the cheapest model, for support reply, long document analysis, full repository review, and structured classification.
Relative cost per workload, indexed to the cheapest model, 2026.

Insights

Monthly spend per developer

The figures below are monthly API spend per developer across 34 engineering teams, each on a single primary model, with the share of that spend from re-sent context and from system prompts and tools.

Primary model Median 75th percentile 90th percentile Share from re-sent context Share from system prompts and tools
GPT-6 Astra $1,180 $2,310 $3,640 61% 24%
Claude Fable 5.1 $2,340 $4,180 $6,920 58% 27%
Gemini 3.1 Pro $395 $810 $1,465 66% 21%
Grok 4.6 $310 $645 $1,190 69% 19%

Insights

Conclusion

A published rate is one input to the bill. On this suite, billed reasoning, prefix caching, and the 200,000-token long-context line move session cost, the cost of 1,000 runs, and monthly spend per developer across a range the rate card alone does not show.

Sources

  1. Frontier Model Token Economics Study. AIR Research Team. September 2026. Houston, TX.
  2. OpenAI, API Pricing. September 2026. San Francisco, California. https://developers.openai.com/api/docs/pricing
  3. Anthropic, Pricing. September 2026. San Francisco, California. https://platform.claude.com/docs/en/about-claude/pricing
  4. Google, Gemini Developer API Pricing. September 2026. Mountain View, California. https://ai.google.dev/gemini-api/docs/pricing
  5. xAI, Models and Pricing. September 2026. Palo Alto, California. https://docs.x.ai/docs/pricing
  6. Artificial Analysis, Independent Analysis of AI Models. September 2026. Sydney, Australia. https://artificialanalysis.ai/models

← All research