THE SYNTHETIC DATA LANDSCAPE: CATEGORIES AND COST DRIVERS
Synthetic training data falls into four categories with distinct GPU profiles. Self-play data: the model generates chains of thought, verifies solutions, and filters correct ones - requires both generation and verification GPUs. Distillation data: a teacher model generates outputs (often with explanations) for the student to train on - requires teacher inference at scale. Prompt-programmed data: hand-crafted prompts with template variables generate diverse examples - cheapest, requires only prompt execution. Augmentation data: existing data transformed via back-translation, paraphrasing, or noise injection - requires small-model inference.
The total GPU cost of synthetic data generation is dominated by the teacher/generator model inference, not the student training. For a 70B teacher generating 100B tokens of synthetic data at $0.10 per million tokens, the generation cost is $10,000. The student training on that data at 7B scale adds $5,000-10,000. The data-generation-to-training cost ratio is typically 1:1 to 3:1 depending on the filtering and verification pipeline overhead.
| Method | GPU Profile | Cost per 1B Tokens | Quality vs Human | Best For |
|---|---|---|---|---|
| Self-play (rejection sampling) | Generator GPU + verifier GPU | $100-500 | 80-95% | Math, reasoning |
| Teacher distillation | Teacher inference (large model) | $50-200 | 85-95% | General purpose |
| Prompt programming | Small model inference | $10-50 | 60-80% | Domain-specific |
| Data augmentation | Small model inference | $5-20 | 70-85% | NLP augmentation |
| Licensed (marketplace) | Zero GPU (purchase) | $50-500 | 95-100% | High-stakes tasks |
| Human annotation | Zero GPU (labor) | $5,000-50,000 | 100% | Fine-tuning, eval |
SELF-PLAY AND REJECTION SAMPLING: THE COMPUTE INTENSIVE APPROACH
Self-play generates multiple candidate responses per prompt and filters for correct ones. For mathematical reasoning, a 70B model generates 64-256 samples per problem, solving 40-70 percent correctly depending on difficulty. At 64 samples per prompt with 512-token responses and 1M prompts, total generation is 64 x 512 x 1M = 33B tokens of raw generation, of which 13-23B tokens are correct solutions. The compute cost: 33B tokens at $0.10/M tokens on H100 = $3,300 for generation, plus a verifier (typically a smaller 7B model or reward model) costing another $300-500.
The rejection sample cost is the correct-to-total ratio. For easy problems (90 percent solve rate), generating 1 correct solution costs approximately 1.1x the cost of a single generation. For hard problems (10 percent solve rate, common in frontier math benchmarks), generating 1 correct solution requires 10 generations, each consuming 500-2,000 tokens of compute. The average cost per correct solution for hard mathematical problems is $0.02-0.10 on H100 - comparable to the per-token training cost for the downstream model.
DeepSeek-R1's self-play took this to an extreme: generating 1M+ reasoning traces at 8K-32K tokens each with an ensemble of verifiers. The total synthetic data generation cost for R1 was estimated at $5-15M in GPU compute - a significant portion of the total training budget. The key insight: self-play data disproportionately generates long, complex trajectories, which are 10-100x more expensive per token than simple demonstrations but contribute disproportionately to reasoning capability.
| Scenario | Samples per Prompt | Avg Correct Rate | Cost per 1M Correct | Cost Ratio | GPU Hours (1M Problems) |
|---|---|---|---|---|---|
| Easy math (grade school) | 4 | 90% | $0.02 | 1.1x single | 50 |
| Medium math (AMC level) | 16 | 60% | $0.06 | 2.5x single | 400 |
| Hard math (IMO level) | 256 | 15% | $0.35 | 16x single | 6,000 |
| Code generation (HumanEval) | 8 | 75% | $0.04 | 1.5x single | 150 |
| Code (competitive prog) | 128 | 20% | $0.25 | 8x single | 4,500 |
| Reasoning (AIME level) | 256 | 10% | $0.60 | 20x single | 12,000 |
DISTILLATION DATA GENERATION: TEACHER COST VERSUS STUDENT VALUE
Generating synthetic data via teacher distillation has a favorable ROI when the teacher is reused. A 405B teacher generating 100B tokens costs approximately $50,000-100,000 in H100 inference compute. This same 100B-token dataset can train multiple student models (7B, 13B, 34B) each costing $10,000-50,000 in training compute. If the dataset enables 1-2 percent accuracy improvement on the student, the value across 10 product deployments could be $500,000+ - a 3-10x ROI on the generation cost.
The amortization factor is critical: teacher-generated synthetic data for alignment (RLHF preferences, safety classifications, style guidelines) retains value across model versions. A single 1B-token preference dataset generated by a 405B teacher costs $500-1,000 but can guide model alignment across 10+ training iterations. The ongoing cost of synthetic data generation is often 5-15 percent of the total training budget, with 80 percent of that cost concentrated in the initial data generation phase before the first model training run.
The counterpoint: synthetic data has a quality ceiling. A 7B student trained on 405B teacher data typically achieves 90-95 percent of the teacher's benchmark performance but cannot surpass the teacher's capability ceiling for factual accuracy. For surpassing the teacher, self-play data (where the student generates and self-verifies) is necessary but 3-10x more expensive per quality-adjusted token.
| Data Type | Teacher Size | Gen Cost 1B Tokens | Student Gain (7B) | Amortized Value | ROI |
|---|---|---|---|---|---|
| General text (SFT) | 70B teacher | $70-150 | +1-3% on MMLU | $5K-15K | 50-100x |
| Math reasoning | 405B teacher | $200-500 | +3-8% on GSM8K | $10K-30K | 20-60x |
| Code explanations | DeepSeek-Coder 236B | $150-300 | +2-5% on HumanEval | $8K-20K | 25-70x |
| Preference data (DPO) | 405B + reward | $500-1,000 | +1-2% win rate | $3K-10K | 3-10x |
| Self-play (same model) | Generator + verifier | $300-800 | +5-15% on hard | Intrinsic | N/A (quality) |
LICENSED DATA: WHEN BUYING BEATS GENERATING
Data marketplaces (Scale AI, Defined.ai, Appen, HumanSignal) charge $50-500 per million tokens for curated training data. At $100/M tokens, 10B tokens of licensed data costs $1M. The equivalent synthetic generation using a 70B teacher costs approximately $1,000-2,000 in GPU compute - 500x cheaper. However, licensed data is human-verified and legally clean (copyright-cleared), whereas synthetic data carries legal risk from model output copyright and potential model collapse from training on generated data.
The break-even analysis favors synthetic data for volume at 100M+ tokens. For small datasets (<10M tokens), the fixed cost of setting up a synthetic generation pipeline ($5,000-20,000 in engineer time, $500-2,000 in prompt development) exceeds the $500-1,000 licensing cost for the same volume. For medium datasets (10M-100M tokens), it's a tie: synthetic generation costs $500-5,000 in GPU compute versus $1,000-10,000 licensing, with the synthetic option adding 1-2 weeks of pipeline development.
For high-stakes domains (medical, legal, financial), the legal clarity of licensed data justifies 5-10x price premium over synthetic alternatives. The current market trend is hybrid: 70-80 percent synthetic data from multiple generators (different model sizes and prompt strategies to increase diversity) plus 20-30 percent licensed human data for quality anchoring and copyright safety. The GPU-to-licensing budget ratio for a typical training run: 15-25 percent GPU (synthetic generation) + 5-10 percent licensing + 65-80 percent training compute.
| Data Volume | Synthetic GPU Cost | Licensed Cost | Quality Difference | Legal Risk | Recommendation |
|---|---|---|---|---|---|
| 1M tokens | $0.10-0.50 (if pipeline exists) | $50-200 | Licensed > Synthetic | Low for licensed | Buy licensed |
| 10M tokens | $1-5 | $500-2,000 | Comparable | Medium synthetic | Tie (depends) |
| 100M tokens | $10-50 | $5,000-20,000 | Comparable | Medium synthetic | Generate |
| 1B tokens | $100-500 | $50,000-200,000 | Synthetic can exceed | Medium synthetic | Generate |
| 10B tokens | $1,000-5,000 | $500K-2M | Synthetic wins ratio | Mitigate via mix | Generate + 10% licensed |
| 100B tokens | $10,000-50,000 | $5M-20M | Synthetic only feasible | Mitigate via mix | Generate + 5% licensed |
QUALITY AND DIVERSITY: THE HIDDEN GPU COST OF SYNTHETIC DATA
Synthetic data suffers from mode collapse: multiple prompts to the same model produce similar outputs, especially at the same temperature. At temperature 0.7, the diversity of a 70B model's outputs across 100 repeated prompts is approximately 40-50 percent lower than human-written text on the same topics. Mitigating mode collapse requires multi-model generation (3-5 different generator models), temperature variation (0.3-1.2), and prompt perturbation - each adding 2-5x to the generation GPU cost.
The hidden cost is deduplication and filtering of synthetic data. A 1B-token synthetic corpus typically contains 30-50 percent near-duplicate content. Deduplication (MinHash, SimHash) on 1B tokens costs $50-100 in CPU compute and 5-10 GPU-hours for embedding-based deduplication. Quality filtering via a smaller reference model adds another 10-20 percent cost. The effective real cost per non-duplicate, high-quality synthetic token is 1.5-2x the raw generation cost.
Training on low-quality synthetic data can permanently degrade model quality - a phenomenon called model collapse. Once a model is trained on synthetic data with systematic errors (e.g., wrong chain-of-thought reasoning steps), subsequent training iterations compound the errors. The mitigation cost is periodic human evaluation: $1,000-5,000 per evaluation cycle for 1,000-5,000 human-rated samples. Evaluation should run every 10-20 training checkpoints, adding 1-2 percent to the total training budget.
B200: MAKING SYNTHETIC DATA ECONOMICALLY FEASIBLE AT SCALE
B200 reduces synthetic data generation cost by 40-50 percent through higher throughput per GPU and reduced KV cache constraints for long generations. For self-play data with 8K-32K token responses, B200's 192 GB memory enables batch sizes 2-4x larger than H100 for the same generation task, directly translating to 2-4x higher throughput for the generator model. A 405B teacher generating synthetic data on B200 achieves 1,200 tokens/s/GPU versus 700 on H100 (FP8), reducing the per-token generation cost by 42 percent.
The combined effect for a large synthetic data pipeline: generating 100B tokens with a 405B teacher costs $58,000 on H100 (64 GPUs, 28 days) versus $25,000 on B200 (32 GPUs, 14 days). The 57 percent cost reduction and 50 percent time reduction make large-scale synthetic data generation practical for mid-sized labs. For the self-play pipeline generating 5B tokens of hard reasoning traces, B200 reduces compute time from 12 days to 5 days on equivalent GPUs. At the organizational level, the hybrid strategy becomes straightforward: use B200 for all synthetic generation, allocate 10-15 percent of GPU budget to data creation, 10-15 percent to filtering and evaluation, and 70-80 percent to training.
