Home / Research / Synthetic Data Statistics
Market research · September 18, 2026
Synthetic Data Statistics: 2026 Report
This report looks at synthetic data use in machine learning as of 2026. Synthetic data means records produced by a model or a simulator rather than observed in the world.
We surveyed 412 organizations running machine learning in production, 61 percent of which had at least one model serving live traffic, spanning financial services, healthcare, retail, industrial, and software. Alongside the survey, we priced synthetic generation against collected and labeled data across seven modalities and tracked accuracy across 1,140 training runs.
Synthetic Data Adoption by Use Case
In the table below, we report the share of surveyed organizations using synthetic data for each purpose, the median share of the data in that workflow that is synthetic, and the driver respondents ranked first.
The Synthetic Data Adoption by Use Case, 2026
| Use Case | Organizations Using | Median Synthetic Share | Top Reported Driver |
|---|---|---|---|
| Software and QA test data | 58% | 27% | Privacy compliance |
| Training set augmentation | 46% | 18% | Edge-case coverage |
| Rare-event and edge-case generation | 41% | 34% | Edge-case coverage |
| Privacy-safe analytics sharing | 36% | 22% | Privacy compliance |
| Fraud and anomaly simulation | 24% | 31% | Class imbalance |
| Computer vision scene generation | 21% | 46% | Cost of collection |
| Agent and assistant evaluation | 19% | 39% | Cost of collection |
Insights
- We found that 63 percent of organizations using synthetic data use it for at least two of the purposes above, and that the ones doing so report materially better outcomes than single-purpose adopters.
- Our data showed the highest synthetic shares in computer vision and agent evaluation, the two areas where collecting the real equivalent is either physically expensive or impossible to arrange on demand.
- We recorded privacy compliance as the top driver in the two highest-adoption use cases and cost as the top driver in the two lowest, which suggests the category is being pulled by regulation more than by budget.
Synthetic Share of Training Data by Modality, 2021 to 2028
Forecasts that a majority of all training data would be synthetic by the middle of the decade have been repeated widely, and they conflate three modalities that have moved at very different speeds. In the table below, we separate them, then plot the same series.
The Synthetic Share of Training Data by Modality, 2026
| Year | Computer Vision | Tabular | Language |
|---|---|---|---|
| 2021 | 6% | 4% | 1% |
| 2022 | 15% | 9% | 3% |
| 2023 | 27% | 15% | 8% |
| 2024 | 38% | 24% | 18% |
| 2025 | 52% | 32% | 27% |
| 2026 | 62% | 41% | 38% |
| 2027 (projected) | 69% | 48% | 49% |
| 2028 (projected) | 74% | 56% | 57% |
Computer vision crossed the halfway mark in 2025 and the other two modalities have not, for reasons that are mechanical rather than cultural. A rendered scene carries its own labels, so a simulator produces perfectly annotated data at the marginal cost of compute, and the gap between rendered and photographed imagery has narrowed to the point where models trained on the former transfer reliably. Tabular generation is constrained by the difficulty of preserving the joint distribution across correlated columns, and language generation is constrained by the fact that the generator and the consumer are increasingly the same class of model, which is the condition under which quality degrades across generations.
Synthetic Generation Cost Against Human Labeling
In the table below, we price 1,000 usable records of each type, generated synthetically and collected and labeled by people, and report the time to generate a million records at the same quality bar.
The Synthetic Generation Cost Against Human Labeling, 2026
| Data Type | Labeled Cost per 1,000 | Synthetic Cost per 1,000 | Cost Reduction | Time to Generate 1M |
|---|---|---|---|---|
| Tabular customer records | $95 | $4 | -96% | 0.8 hours |
| Image classification labels | $68 | $9 | -87% | 3.5 hours |
| Named entity annotations | $135 | $12 | -91% | 2.1 hours |
| Object detection boxes | $420 | $31 | -93% | 6.2 hours |
| Semantic segmentation masks | $1,150 | $74 | -94% | 11.4 hours |
| Audio transcription (per 1,000 minutes) | $1,480 | $210 | -86% | 7.9 hours |
| Medical imaging items | $2,650 | $198 | -93% | 14.6 hours |
The cost reduction is real, and most of these programs still get funded for reasons beyond unit cost. Generation cost per record falls to near zero, and the work moves to validation: deciding whether a generated distribution matches the real one closely enough to train on, and proving it to whoever signs off on the model. In our sample, validation consumed 44 percent of the total hours on a synthetic data project against 11 percent on a human labeling project, which closes roughly a third of the headline saving. The saving that survives is still large, and it is largest in exactly the categories where human labeling is slowest, which is why segmentation and medical imaging show the highest adoption of generation among teams that also report the strictest review processes.
Model Performance by Training Data Mix
In the table below, we report what happened to model accuracy across 1,140 training runs in our sample, indexed against an identical model trained only on real data, along with the number of successive generations of self-training before we measured a statistically significant degradation.
The Model Performance by Training Data Mix, 2026
| Training Mix | Accuracy vs Real-Only | Runs That Beat Baseline | Generations to Degradation |
|---|---|---|---|
| 100% real (baseline) | 100.0 | Not applicable | Not observed |
| 75% real, 25% synthetic | 103.1 | 71% | Not observed |
| 50% real, 50% synthetic | 101.4 | 58% | 9 |
| 25% real, 75% synthetic | 96.8 | 34% | 6 |
| 100% synthetic | 88.2 | 12% | 4 |
Insights
- We found the optimum at roughly a quarter synthetic, where the generated examples add coverage of cases the real set underrepresents without displacing the distribution the model needs to learn.
- Our data showed accuracy falling below the real-only baseline once synthetic content passed 60 percent of the training set, with the decline steepening sharply after that point.
- We observed degradation onset within four generations of pure self-training, and within nine when a real data floor of half the set was maintained, which makes the floor the single most protective design choice available.
Synthetic Data Market Size and Validation Practice
In the table below, we set market size against the share of buyers who could describe a documented process for validating that generated data matches the distribution it is standing in for.
The Synthetic Data Market Size and Validation Practice, 2026
| Year | Market Size | Year-Over-Year Growth | Buyers With a Validation Framework |
|---|---|---|---|
| 2023 | $310M | Not measured | 12% |
| 2024 | $452M | +46% | 15% |
| 2025 | $617M | +37% | 19% |
| 2026 | $814M | +32% | 24% |
| 2030 (projected) | $2.9B | +37% annualized | 46% |
A market growing at better than 30 percent a year in which three quarters of buyers cannot describe how they check the product is a market with a correction in it. What goes wrong is quiet. A generated dataset that is statistically plausible but misses a tail the real data contained produces a model that performs well in evaluation and badly on the cases that matter, and the gap shows up in production rather than in testing. Teams in our sample took a median of eleven days to detect a degradation traceable to generated training data, against two days for one traceable to a pipeline error, because the former looks like ordinary drift until someone checks the source. The providers likely to hold their growth through the next two years are the ones selling validation alongside generation, which today is a minority of them.
Requesting a Copy of This Report
If you would like a PDF copy of this report, or the survey instrument and run-level data behind it, you can reach out here.
Sources
- Synthetic Data Study, Independent Research Team, September 2026, New York, New York.
- Synthetic Data Adoption Rises, but Risks Remain, No Jitter, July 2026, New York, New York. https://www.nojitter.com/data-management/synthetic-data-adoption-rises-but-risks-remain
- Synthetic Data Generation Market Size and Trends, Precedence Research, 2026, Ottawa, Ontario. https://www.precedenceresearch.com/synthetic-data-generation-market
- Synthetic Data Hits $791M: The Proof Is Missing, Luiz Neto, 2026, Sao Paulo, Brazil. https://www.luizneto.ai/synthetic-data-enterprise-2026/
- How Much Does Data Labeling Cost, GigaBPO, 2026, Los Angeles, California. https://gigabpo.com/how-much-does-data-labeling-cost/
- Synthetic Data vs Real Data: When to Use Each for ML Training, Label Your Data, 2026, Kyiv, Ukraine. https://labelyourdata.com/articles/synthetic-data-vs-real-data
