All essays
BenchmarkCOMPARISONFEB 2026

Synthetic Data Generation for Training: GPU Cost of Creating Training Data vs Buying It

Cost-benefit analysis of synthetic data generation versus licensed data acquisition for AI training. GPU compute costs for self-play, model distillation, and prompt programming data generation compared to data marketplace pricing. Benchmarks for code, math, reasoning, and multilingual data at billion-token scale.

01

THE SYNTHETIC DATA LANDSCAPE: CATEGORIES AND COST DRIVERS

Synthetic training data falls into four categories with distinct GPU profiles. Self-play data: the model generates chains of thought, verifies solutions, and filters correct ones - requires both generation and verification GPUs. Distillation data: a teacher model generates outputs (often with explanations) for the student to train on - requires teacher inference at scale. Prompt-programmed data: hand-crafted prompts with template variables generate diverse examples - cheapest, requires only prompt execution. Augmentation data: existing data transformed via back-translation, paraphrasing, or noise injection - requires small-model inference.

The total GPU cost of synthetic data generation is dominated by the teacher/generator model inference, not the student training. For a 70B teacher generating 100B tokens of synthetic data at $0.10 per million tokens, the generation cost is $10,000. The student training on that data at 7B scale adds $5,000-10,000. The data-generation-to-training cost ratio is typically 1:1 to 3:1 depending on the filtering and verification pipeline overhead.

MethodGPU ProfileCost per 1B TokensQuality vs HumanBest For
Self-play (rejection sampling)Generator GPU + verifier GPU$100-50080-95%Math, reasoning
Teacher distillationTeacher inference (large model)$50-20085-95%General purpose
Prompt programmingSmall model inference$10-5060-80%Domain-specific
Data augmentationSmall model inference$5-2070-85%NLP augmentation
Licensed (marketplace)Zero GPU (purchase)$50-50095-100%High-stakes tasks
Human annotationZero GPU (labor)$5,000-50,000100%Fine-tuning, eval
02

SELF-PLAY AND REJECTION SAMPLING: THE COMPUTE INTENSIVE APPROACH

Self-play generates multiple candidate responses per prompt and filters for correct ones. For mathematical reasoning, a 70B model generates 64-256 samples per problem, solving 40-70 percent correctly depending on difficulty. At 64 samples per prompt with 512-token responses and 1M prompts, total generation is 64 x 512 x 1M = 33B tokens of raw generation, of which 13-23B tokens are correct solutions. The compute cost: 33B tokens at $0.10/M tokens on H100 = $3,300 for generation, plus a verifier (typically a smaller 7B model or reward model) costing another $300-500.

The rejection sample cost is the correct-to-total ratio. For easy problems (90 percent solve rate), generating 1 correct solution costs approximately 1.1x the cost of a single generation. For hard problems (10 percent solve rate, common in frontier math benchmarks), generating 1 correct solution requires 10 generations, each consuming 500-2,000 tokens of compute. The average cost per correct solution for hard mathematical problems is $0.02-0.10 on H100 - comparable to the per-token training cost for the downstream model.

DeepSeek-R1's self-play took this to an extreme: generating 1M+ reasoning traces at 8K-32K tokens each with an ensemble of verifiers. The total synthetic data generation cost for R1 was estimated at $5-15M in GPU compute - a significant portion of the total training budget. The key insight: self-play data disproportionately generates long, complex trajectories, which are 10-100x more expensive per token than simple demonstrations but contribute disproportionately to reasoning capability.

ScenarioSamples per PromptAvg Correct RateCost per 1M CorrectCost RatioGPU Hours (1M Problems)
Easy math (grade school)490%$0.021.1x single50
Medium math (AMC level)1660%$0.062.5x single400
Hard math (IMO level)25615%$0.3516x single6,000
Code generation (HumanEval)875%$0.041.5x single150
Code (competitive prog)12820%$0.258x single4,500
Reasoning (AIME level)25610%$0.6020x single12,000
03

DISTILLATION DATA GENERATION: TEACHER COST VERSUS STUDENT VALUE

Generating synthetic data via teacher distillation has a favorable ROI when the teacher is reused. A 405B teacher generating 100B tokens costs approximately $50,000-100,000 in H100 inference compute. This same 100B-token dataset can train multiple student models (7B, 13B, 34B) each costing $10,000-50,000 in training compute. If the dataset enables 1-2 percent accuracy improvement on the student, the value across 10 product deployments could be $500,000+ - a 3-10x ROI on the generation cost.

The amortization factor is critical: teacher-generated synthetic data for alignment (RLHF preferences, safety classifications, style guidelines) retains value across model versions. A single 1B-token preference dataset generated by a 405B teacher costs $500-1,000 but can guide model alignment across 10+ training iterations. The ongoing cost of synthetic data generation is often 5-15 percent of the total training budget, with 80 percent of that cost concentrated in the initial data generation phase before the first model training run.

The counterpoint: synthetic data has a quality ceiling. A 7B student trained on 405B teacher data typically achieves 90-95 percent of the teacher's benchmark performance but cannot surpass the teacher's capability ceiling for factual accuracy. For surpassing the teacher, self-play data (where the student generates and self-verifies) is necessary but 3-10x more expensive per quality-adjusted token.

Data TypeTeacher SizeGen Cost 1B TokensStudent Gain (7B)Amortized ValueROI
General text (SFT)70B teacher$70-150+1-3% on MMLU$5K-15K50-100x
Math reasoning405B teacher$200-500+3-8% on GSM8K$10K-30K20-60x
Code explanationsDeepSeek-Coder 236B$150-300+2-5% on HumanEval$8K-20K25-70x
Preference data (DPO)405B + reward$500-1,000+1-2% win rate$3K-10K3-10x
Self-play (same model)Generator + verifier$300-800+5-15% on hardIntrinsicN/A (quality)
04

LICENSED DATA: WHEN BUYING BEATS GENERATING

Data marketplaces (Scale AI, Defined.ai, Appen, HumanSignal) charge $50-500 per million tokens for curated training data. At $100/M tokens, 10B tokens of licensed data costs $1M. The equivalent synthetic generation using a 70B teacher costs approximately $1,000-2,000 in GPU compute - 500x cheaper. However, licensed data is human-verified and legally clean (copyright-cleared), whereas synthetic data carries legal risk from model output copyright and potential model collapse from training on generated data.

The break-even analysis favors synthetic data for volume at 100M+ tokens. For small datasets (<10M tokens), the fixed cost of setting up a synthetic generation pipeline ($5,000-20,000 in engineer time, $500-2,000 in prompt development) exceeds the $500-1,000 licensing cost for the same volume. For medium datasets (10M-100M tokens), it's a tie: synthetic generation costs $500-5,000 in GPU compute versus $1,000-10,000 licensing, with the synthetic option adding 1-2 weeks of pipeline development.

For high-stakes domains (medical, legal, financial), the legal clarity of licensed data justifies 5-10x price premium over synthetic alternatives. The current market trend is hybrid: 70-80 percent synthetic data from multiple generators (different model sizes and prompt strategies to increase diversity) plus 20-30 percent licensed human data for quality anchoring and copyright safety. The GPU-to-licensing budget ratio for a typical training run: 15-25 percent GPU (synthetic generation) + 5-10 percent licensing + 65-80 percent training compute.

Data VolumeSynthetic GPU CostLicensed CostQuality DifferenceLegal RiskRecommendation
1M tokens$0.10-0.50 (if pipeline exists)$50-200Licensed > SyntheticLow for licensedBuy licensed
10M tokens$1-5$500-2,000ComparableMedium syntheticTie (depends)
100M tokens$10-50$5,000-20,000ComparableMedium syntheticGenerate
1B tokens$100-500$50,000-200,000Synthetic can exceedMedium syntheticGenerate
10B tokens$1,000-5,000$500K-2MSynthetic wins ratioMitigate via mixGenerate + 10% licensed
100B tokens$10,000-50,000$5M-20MSynthetic only feasibleMitigate via mixGenerate + 5% licensed
05

QUALITY AND DIVERSITY: THE HIDDEN GPU COST OF SYNTHETIC DATA

Synthetic data suffers from mode collapse: multiple prompts to the same model produce similar outputs, especially at the same temperature. At temperature 0.7, the diversity of a 70B model's outputs across 100 repeated prompts is approximately 40-50 percent lower than human-written text on the same topics. Mitigating mode collapse requires multi-model generation (3-5 different generator models), temperature variation (0.3-1.2), and prompt perturbation - each adding 2-5x to the generation GPU cost.

The hidden cost is deduplication and filtering of synthetic data. A 1B-token synthetic corpus typically contains 30-50 percent near-duplicate content. Deduplication (MinHash, SimHash) on 1B tokens costs $50-100 in CPU compute and 5-10 GPU-hours for embedding-based deduplication. Quality filtering via a smaller reference model adds another 10-20 percent cost. The effective real cost per non-duplicate, high-quality synthetic token is 1.5-2x the raw generation cost.

Training on low-quality synthetic data can permanently degrade model quality - a phenomenon called model collapse. Once a model is trained on synthetic data with systematic errors (e.g., wrong chain-of-thought reasoning steps), subsequent training iterations compound the errors. The mitigation cost is periodic human evaluation: $1,000-5,000 per evaluation cycle for 1,000-5,000 human-rated samples. Evaluation should run every 10-20 training checkpoints, adding 1-2 percent to the total training budget.

06

B200: MAKING SYNTHETIC DATA ECONOMICALLY FEASIBLE AT SCALE

B200 reduces synthetic data generation cost by 40-50 percent through higher throughput per GPU and reduced KV cache constraints for long generations. For self-play data with 8K-32K token responses, B200's 192 GB memory enables batch sizes 2-4x larger than H100 for the same generation task, directly translating to 2-4x higher throughput for the generator model. A 405B teacher generating synthetic data on B200 achieves 1,200 tokens/s/GPU versus 700 on H100 (FP8), reducing the per-token generation cost by 42 percent.

The combined effect for a large synthetic data pipeline: generating 100B tokens with a 405B teacher costs $58,000 on H100 (64 GPUs, 28 days) versus $25,000 on B200 (32 GPUs, 14 days). The 57 percent cost reduction and 50 percent time reduction make large-scale synthetic data generation practical for mid-sized labs. For the self-play pipeline generating 5B tokens of hard reasoning traces, B200 reduces compute time from 12 days to 5 days on equivalent GPUs. At the organizational level, the hybrid strategy becomes straightforward: use B200 for all synthetic generation, allocate 10-15 percent of GPU budget to data creation, 10-15 percent to filtering and evaluation, and 70-80 percent to training.

Filed under
Synthetic Data GenerationGPU Data GenerationSelf Play TrainingData Distillation GPULicensed Training DataData MarketplaceTraining Data Economics