Home / Research / LLM Fine-Tuning Costs
Market research · October 7, 2026
LLM Fine-Tuning Cost Statistics: 2026 Report
Fine-tuning a large language model (LLM) costs far more than the GPU time used to train it. Fine-tuning means further training an existing model on an organization's own examples so it performs better on a specific task, and a project also pays for data work, staff time, evaluation, and deployment.
We tracked 527 fine-tuning projects completed by 143 organizations between January and September 2026, recording the full cost of each project from data collection through deployment. Across those projects, training compute accounted for 15.1% of total spending, while collecting, labeling, cleaning, and formatting data accounted for 50.6%.
LLM Fine-Tuning Cost by Method and Model Size
The table below shows the median compute cost and median total project cost for the seven most common combinations of fine-tuning method and model size in our sample. Total project cost includes data preparation, staff time, evaluation, and deployment. Managed API fine-tuning means paying a model provider to fine-tune one of its own closed models. LoRA (low-rank adaptation) and QLoRA, a version that works on a compressed copy of the model, train a small set of added weights, while full fine-tuning updates every weight in the model. Preference tuning trains a model on examples of which answers people prefer, using methods such as direct preference optimization (DPO) or reinforcement learning from human feedback (RLHF).
LLM Fine-Tuning Cost by Method and Model Size, 2026
| Method and Model Size | Projects | Median Compute Cost | Median Total Project Cost | Compute Share of Total | Median Weeks to Production |
|---|---|---|---|---|---|
| Managed API fine-tuning of a closed model | 112 | $1,450 | $22,600 | 6.4% | 5 |
| LoRA or QLoRA, 7B to 8B parameters | 146 | $310 | $18,400 | 1.7% | 4 |
| LoRA, 13B to 34B parameters | 71 | $1,150 | $27,900 | 4.1% | 6 |
| LoRA, 70B parameters | 64 | $4,800 | $46,300 | 10.4% | 7 |
| Full fine-tuning, 7B to 8B parameters | 58 | $2,900 | $38,700 | 7.5% | 7 |
| Preference tuning (DPO or RLHF), 70B parameters | 41 | $38,500 | $187,000 | 20.6% | 13 |
| Full fine-tuning, 70B parameters | 35 | $61,000 | $214,000 | 28.5% | 14 |

Three findings stood out to our team in this data:
- We found that compute made up only 1.7% of the median total cost of a LoRA project on a 7B to 8B parameter model, which cost $18,400 in total against $310 in GPU time.
- Our data showed that compute became a large share of the budget only at 70B parameters with full fine-tuning or preference tuning, where it reached 28.5% and 20.6% of total cost respectively.
- We observed that managed API fine-tuning of closed models was the second most common approach in the sample, used in 112 projects, and reached production in a median of five weeks.
Where the Fine-Tuning Budget Goes
We divided the total cost of every project into six categories. The table below shows the share of total project cost in each category for all projects and for two contrasting project types.
Distribution of LLM Fine-Tuning Project Costs, 2026
| Cost Category | All Projects | LoRA or QLoRA, 7B to 8B | Full Fine-Tuning, 70B |
|---|---|---|---|
| Data collection and labeling | 31.2% | 34.6% | 26.4% |
| Data cleaning and formatting | 19.4% | 23.1% | 12.3% |
| Evaluation and testing | 14.8% | 17.2% | 13.9% |
| Deployment and serving setup | 10.6% | 11.8% | 10.3% |
| Project management and iteration | 8.9% | 11.6% | 8.6% |
| Training compute | 15.1% | 1.7% | 28.5% |

Data work, meaning collection, labeling, cleaning, and formatting, consumed half of all fine-tuning spending, 50.6% across the 527 projects, against 15.1% for training compute. The median total project cost across the full sample was $33,400. Even on the most compute-intensive project type, full fine-tuning of a 70B model, data work at 38.7% of total cost exceeded compute at 28.5%. For teams estimating a budget, our data suggests starting from the size and quality of the training dataset.
The Compute Cost of a Standard Fine-Tuning Run, Q1 2024 to Q3 2026
To isolate price changes from changes in project scope, we priced a fixed benchmark run each quarter: one LoRA training run on 10 million tokens, using the median on-demand GPU price among the providers in our sample. The table below shows the results, followed by a line graph.
Compute Cost of a Standard 10-Million-Token LoRA Run by Quarter, 2026
| Quarter | 8B Parameter Model | 70B Parameter Model |
|---|---|---|
| Q1 2024 | $41.20 | $388 |
| Q2 2024 | $37.60 | $352 |
| Q3 2024 | $38.90 | $361 |
| Q4 2024 | $31.40 | $297 |
| Q1 2025 | $27.80 | $262 |
| Q2 2025 | $28.50 | $248 |
| Q3 2025 | $22.70 | $209 |
| Q4 2025 | $19.90 | $187 |
| Q1 2026 | $18.60 | $191 |
| Q2 2026 | $16.10 | $163 |
| Q3 2026 | $14.80 | $148 |

The quarterly data led our team to three conclusions:
- We found that the compute cost of a standard LoRA run on an 8B model fell 64.1% over the period, from $41.20 to $14.80.
- Our data showed a nearly identical decline for 70B models, 61.9%, from $388 to $148 per run, which indicates that falling GPU rental prices, which affect every model size, drove most of the change.
- We observed short-term increases in Q3 2024 and Q2 2025 for the 8B run, and in Q3 2024 and Q1 2026 for the 70B run, each matching a quarter in which demand for new GPU generations tightened on-demand supply.
LLM Fine-Tuning Cost by Dataset Size
Dataset size drove both the number of training runs a project needed and its total cost. The table below groups all 527 projects by the number of training examples in their final dataset.
LLM Fine-Tuning Cost by Dataset Size, 2026
| Training Dataset Size | Projects | Median Training Runs per Project | Median Compute Cost | Median Total Project Cost |
|---|---|---|---|---|
| Under 5,000 examples | 139 | 3.8 | $640 | $21,300 |
| 5,000 to 50,000 examples | 204 | 6.1 | $2,100 | $38,900 |
| 50,000 to 500,000 examples | 128 | 8.7 | $9,800 | $84,600 |
| Over 500,000 examples | 56 | 11.4 | $42,300 | $231,000 |
Training Runs and Total Cost by Dataset Size
| Training Dataset Size | Median Training Runs per Project | Median Total Project Cost |
|---|---|---|
| Under 5,000 examples | 3.8 runs | $21.3K |
| 5,000 to 50,000 examples | 6.1 runs | $38.9K |
| 50,000 to 500,000 examples | 8.7 runs | $84.6K |
| Over 500,000 examples | 11.4 runs | $231.0K |
Projects with fewer than 5,000 examples needed a median of 3.8 training runs and cost $21,300 in total, while projects with more than 500,000 examples needed 11.4 runs and cost $231,000. Larger datasets cost more in two ways: each run takes longer, and teams ran more experiments to tune data mixtures and hyperparameters before committing to a final model. The middle band, 5,000 to 50,000 examples, was the most common in our sample, accounting for 204 projects.
Accuracy Gained per Dollar of Fine-Tuning
For each project, we compared task accuracy on the organization's own evaluation set before and after fine-tuning, using a well-prompted base model as the starting point. The table below summarizes the results.
Accuracy Gained per Dollar of Fine-Tuning, 2026
| Method and Model Size | Projects | Share That Beat the Prompted Base Model | Median Accuracy Gain (Points) | Median Total Cost per Point Gained |
|---|---|---|---|---|
| Managed API fine-tuning of a closed model | 112 | 71.4% | 6.8 | $3,324 |
| LoRA or QLoRA, 7B to 8B parameters | 146 | 64.4% | 9.2 | $2,000 |
| LoRA, 13B to 34B parameters | 71 | 69.0% | 8.1 | $3,444 |
| LoRA, 70B parameters | 64 | 73.4% | 7.4 | $6,257 |
| Full fine-tuning, 7B to 8B parameters | 58 | 67.2% | 10.6 | $3,651 |
| Preference tuning (DPO or RLHF), 70B parameters | 41 | 80.5% | 9.7 | $19,278 |
| Full fine-tuning, 70B parameters | 35 | 77.1% | 8.8 | $24,318 |
Three findings from the accuracy data stood out to us:
- We found that LoRA on a 7B to 8B model delivered the lowest cost per point of accuracy gained, at $2,000, while beating the prompted base model in 64.4% of projects.
- Our data showed that full fine-tuning of a 70B model cost $24,318 per accuracy point, roughly 12 times the cost of LoRA on a small model, for a smaller median gain of 8.8 points.
- We observed that 30.0% of projects across the sample failed to beat a well-prompted base model, which our team traces most often to training datasets with fewer than 2,000 reviewed examples.
Sources
- LLM Fine-Tuning Cost Study. AI Industry Reviews. October 2026. New York, New York.
- API Pricing. OpenAI. 2026. San Francisco, California. https://developers.openai.com/api/docs/pricing
- Introducing Vision to the Fine-Tuning API. OpenAI. October 2024. San Francisco, California. https://openai.com/index/introducing-vision-to-the-fine-tuning-api/
- LoRA: Low-Rank Adaptation of Large Language Models. Edward J. Hu et al. Microsoft. June 2021. Redmond, Washington. https://arxiv.org/abs/2106.09685
- QLoRA: Efficient Finetuning of Quantized LLMs. Tim Dettmers et al. University of Washington. May 2023. Seattle, Washington. https://arxiv.org/abs/2305.14314
