Home / Research / Synthetic Data Statistics

Market research · September 18, 2026

Synthetic Data Statistics: 2026 Report

This report looks at synthetic data use in machine learning as of 2026. Synthetic data means records produced by a model or a simulator rather than observed in the world.

We surveyed 412 organizations running machine learning in production, 61 percent of which had at least one model serving live traffic, spanning financial services, healthcare, retail, industrial, and software. Alongside the survey, we priced synthetic generation against collected and labeled data across seven modalities and tracked accuracy across 1,140 training runs.

Synthetic Data Adoption by Use Case

In the table below, we report the share of surveyed organizations using synthetic data for each purpose, the median share of the data in that workflow that is synthetic, and the driver respondents ranked first.

The Synthetic Data Adoption by Use Case, 2026

Use CaseOrganizations UsingMedian Synthetic ShareTop Reported Driver
Software and QA test data58%27%Privacy compliance
Training set augmentation46%18%Edge-case coverage
Rare-event and edge-case generation41%34%Edge-case coverage
Privacy-safe analytics sharing36%22%Privacy compliance
Fraud and anomaly simulation24%31%Class imbalance
Computer vision scene generation21%46%Cost of collection
Agent and assistant evaluation19%39%Cost of collection

Insights

Synthetic Share of Training Data by Modality, 2021 to 2028

Forecasts that a majority of all training data would be synthetic by the middle of the decade have been repeated widely, and they conflate three modalities that have moved at very different speeds. In the table below, we separate them, then plot the same series.

The Synthetic Share of Training Data by Modality, 2026

YearComputer VisionTabularLanguage
20216%4%1%
202215%9%3%
202327%15%8%
202438%24%18%
202552%32%27%
202662%41%38%
2027 (projected)69%48%49%
2028 (projected)74%56%57%
Line chart of synthetic share of training data by modality from 2021 to 2028 for Computer Vision, Tabular, and Language, with 2027-2028 projected.
Synthetic Share of Training Data by Modality

Computer vision crossed the halfway mark in 2025 and the other two modalities have not, for reasons that are mechanical rather than cultural. A rendered scene carries its own labels, so a simulator produces perfectly annotated data at the marginal cost of compute, and the gap between rendered and photographed imagery has narrowed to the point where models trained on the former transfer reliably. Tabular generation is constrained by the difficulty of preserving the joint distribution across correlated columns, and language generation is constrained by the fact that the generator and the consumer are increasingly the same class of model, which is the condition under which quality degrades across generations.

Synthetic Generation Cost Against Human Labeling

In the table below, we price 1,000 usable records of each type, generated synthetically and collected and labeled by people, and report the time to generate a million records at the same quality bar.

The Synthetic Generation Cost Against Human Labeling, 2026

Data TypeLabeled Cost per 1,000Synthetic Cost per 1,000Cost ReductionTime to Generate 1M
Tabular customer records$95$4-96%0.8 hours
Image classification labels$68$9-87%3.5 hours
Named entity annotations$135$12-91%2.1 hours
Object detection boxes$420$31-93%6.2 hours
Semantic segmentation masks$1,150$74-94%11.4 hours
Audio transcription (per 1,000 minutes)$1,480$210-86%7.9 hours
Medical imaging items$2,650$198-93%14.6 hours

The cost reduction is real, and most of these programs still get funded for reasons beyond unit cost. Generation cost per record falls to near zero, and the work moves to validation: deciding whether a generated distribution matches the real one closely enough to train on, and proving it to whoever signs off on the model. In our sample, validation consumed 44 percent of the total hours on a synthetic data project against 11 percent on a human labeling project, which closes roughly a third of the headline saving. The saving that survives is still large, and it is largest in exactly the categories where human labeling is slowest, which is why segmentation and medical imaging show the highest adoption of generation among teams that also report the strictest review processes.

Model Performance by Training Data Mix

In the table below, we report what happened to model accuracy across 1,140 training runs in our sample, indexed against an identical model trained only on real data, along with the number of successive generations of self-training before we measured a statistically significant degradation.

The Model Performance by Training Data Mix, 2026

Training MixAccuracy vs Real-OnlyRuns That Beat BaselineGenerations to Degradation
100% real (baseline)100.0Not applicableNot observed
75% real, 25% synthetic103.171%Not observed
50% real, 50% synthetic101.458%9
25% real, 75% synthetic96.834%6
100% synthetic88.212%4

Insights

Synthetic Data Market Size and Validation Practice

In the table below, we set market size against the share of buyers who could describe a documented process for validating that generated data matches the distribution it is standing in for.

The Synthetic Data Market Size and Validation Practice, 2026

YearMarket SizeYear-Over-Year GrowthBuyers With a Validation Framework
2023$310MNot measured12%
2024$452M+46%15%
2025$617M+37%19%
2026$814M+32%24%
2030 (projected)$2.9B+37% annualized46%

A market growing at better than 30 percent a year in which three quarters of buyers cannot describe how they check the product is a market with a correction in it. What goes wrong is quiet. A generated dataset that is statistically plausible but misses a tail the real data contained produces a model that performs well in evaluation and badly on the cases that matter, and the gap shows up in production rather than in testing. Teams in our sample took a median of eleven days to detect a degradation traceable to generated training data, against two days for one traceable to a pipeline error, because the former looks like ordinary drift until someone checks the source. The providers likely to hold their growth through the next two years are the ones selling validation alongside generation, which today is a minority of them.

Requesting a Copy of This Report

If you would like a PDF copy of this report, or the survey instrument and run-level data behind it, you can reach out here.

Sources