The 2026 Training Cost Landscape
Foundation model training costs in 2026 are dominated by GPU compute, but the secondary cost categories -- network egress, storage, power, and data preparation -- have grown in relative importance as GPU prices have fallen. Total training cost for a 70B-parameter dense model on 2,048 H200 GPUs for 30 days is approximately $4.5M at Q2 2026 spot rates. GPU compute accounts for 72% of the total, down from 85% in 2024, as network and storage costs have risen with larger datasets and more frequent checkpointing.
The cost breakdown varies dramatically by model architecture. Mixture-of-experts models (like DeepSeek V3 and Qwen 3-235B) require 40-50% fewer FLOPs per parameter than dense models, but their memory access patterns increase storage I/O costs by 25-35%. Long-context training (128k+ tokens) adds roughly 15% to training cost through increased attention computation and larger checkpoint sizes. Multi-modal training adds data preprocessing and storage costs that can reach 10-15% of total training budget.
GPU Compute: The Dominant Cost Component
GPU compute cost is the product of GPU count, training duration, and per-GPU-hour rate. For a single training run of a 70B parameter dense model using 2,048 H200 GPUs for 30 days at $3.10/GPU/hr spot rate, GPU compute alone costs $4.27M. The per-GPU-hour rate is the largest variable: reserved contracts on H200 run $2.40-2.80/GPU/hr, on-demand spot rates run $2.90-3.50/GPU/hr, and premium providers with guaranteed availability charge $3.80-4.50/GPU/hr.
GPU utilization during training is never 100%. The average H200 cluster achieves 65-75% of theoretical peak FLOP utilization for training, depending on model parallelism strategy and data loading efficiency. At 70% utilization, effective compute cost per useful FLOP is $4.43/GPU-hr, 43% higher than the nominal $3.10/GPU-hr. Blackwell GPUs (B200, B300) achieve 75-82% utilization on the same workloads, narrowing the effective cost gap with Hopper despite higher nominal rates.
| Model Size | GPUs | Duration | GPU Rate | Total GPU Cost |
|---|---|---|---|---|
| 7B dense | 64 H200 | 5 days | $3.10/hr | $23,808 |
| 13B dense | 128 H200 | 10 days | $3.10/hr | $95,232 |
| 70B dense | 2,048 H200 | 30 days | $3.10/hr | $4,271,040 |
| 405B dense | 8,192 H200 | 60 days | $3.10/hr | $36,372,480 |
| 235B MoE | 4,096 H200 | 35 days | $3.10/hr | $10,657,920 |
| 1T+ MoE | 16,384 B300 | 90 days | $5.50/hr | $194,803,200 |
Network Egress and Interconnect Costs
Network egress costs arise from three sources during training: dataset loading (transferring training data from object storage to GPU nodes), checkpoint save/load (transferring model weights and optimizer states to persistent storage), and inter-node communication (all-reduce, all-to-all across the fabric). The first two carry per-byte egress charges from the storage provider. The third is included in the GPU compute rate for most providers but becomes a cost factor when using cross-region or cross-provider cluster configurations.
Dataset loading for a 70B model trained on 15 trillion tokens with a 1:1 raw-to-unique-token ratio requires transferring approximately 40 TB of training data. At $0.01/GB egress from object storage, this adds $410 per training run. Checkpointing is more significant: each checkpoint for a 70B model in BF16 requires 280 GB (weights) plus 420 GB (optimizer states) for FP32 Adam. With checkpoints saved every 2 hours for 30 days (360 checkpoints), total checkpoint storage I/O is 252 TB. At $0.02/GB write cost for parallel filesystems, checkpoint write costs reach $5,040 per training run.
Storage Tiering: Hot, Warm, and Cold Costs
Training infrastructure requires three storage tiers with different cost profiles. Hot storage (parallel filesystem like WekaFS, GPUDirect Storage, or Lustre) is used for active training data and checkpoints. It costs $0.08-0.15/GB/month with 10-50 GB/s throughput. Warm storage (NFS or S3-compatible object storage) holds dataset archives and intermediate checkpoints at $0.01-0.03/GB/month. Cold storage (archival S3, Glacier) holds completed model versions and training logs at $0.001-0.005/GB/month.
For a 70B training run, storage cost breakdown: 50 TB hot filesystem at $0.10/GB/month for 2 months = $10,000; 200 TB warm object storage at $0.02/GB/month for 6 months = $24,000; and 50 TB cold archival storage at $0.003/GB/month ongoing = $150/month. The total storage cost for the training run lifecycle is approximately $34,000 plus $150/month ongoing. Hot storage dominates the short-term cost, but warm storage accumulation over multiple training runs grows into the largest long-term line item.
Power and Cooling: The PUE Tax
Power and cooling costs are typically included in the GPU rental rate, but the allocation varies significantly by provider. GPU providers with low PUE (Power Usage Effectiveness) data centers include less power overhead in their rates. A provider with PUE 1.15 includes 15% additional power cost beyond the GPU's TDP. A provider with PUE 1.40 includes 40% overhead. At $0.12/kWh average industrial electricity cost, the power component of an H200 GPU running at 700W TDP is approximately $0.084/GPU-hr for the GPU alone, plus PUE overhead.
The PUE impact on total training cost is meaningful. For the 2,048-H200, 30-day training run at PUE 1.15, power adds $124,000. At PUE 1.40, power adds $151,000 -- a $27,000 difference driven purely by data center efficiency. Liquid-cooled facilities (PUE 1.05-1.15) command a premium in GPU pricing, typically $0.15-0.30/GPU-hr above air-cooled equivalents, but the power savings offset the premium for workloads lasting longer than 60 days. For shorter training runs, air-cooled GPUs at lower rates are the cost-optimal choice.
Data Processing and Curation Costs
Data preparation costs are frequently omitted from training cost breakdowns but represent 5-12% of total project budget. For a 15-trillion-token dataset, data processing includes deduplication (MinHash, Bloom filters), quality filtering (perplexity scoring, classifier-based filtering), and tokenization. At $0.50/GB processed (including compute and storage), data processing for 40 TB of raw data costs $20,000. Data synthesis and augmentation add another $10,000-50,000 depending on the complexity of the synthetic pipeline.
The hidden cost is repeated processing: teams typically iterate on the dataset 3-5 times per training run, adjusting tokenizer, data mix, or filtering thresholds. Each iteration re-processes the full dataset. A team that does 5 iterations at $20,000 per iteration has spent $100,000 on data processing before the first training job starts. This cost is often blamed on GPU pricing when the real culprit is undisciplined data iteration without proper versioning and caching.
Experimentation and Ablation Run Costs
The published training cost for a foundation model represents only the final production run. Most teams run 20-50 ablation studies per model release. Small-scale ablations (7B models on 64 GPUs for 2-3 days) cost $20,000-30,000 each. Larger ablations (70B models on 512-1,024 GPUs for 5-7 days) cost $200,000-400,000 each. A typical foundation model release involves $3-8M in ablation and experimentation costs before the final training run.
Ablation experiments fall into four categories: architecture decisions (MoE routing, attention variant, normalization placement), data mix optimization (internet data ratio, code data ratio, synthetic data percentage), hyperparameter sweeps (learning rate, batch size, weight decay schedule), and training stability tests (gradient clipping, loss scaling, checkpoint validation). Each category typically requires 5-15 experiments. The cumulative cost of failed or superseded experiments is a significant but often unreported component of total training expenditure.
| Cost Category | 70B Dense Model | 235B MoE Model | 1T+ Frontier Model |
|---|---|---|---|
| GPU Compute | $4,271,040 | $10,657,920 | $194,803,200 |
| Network Egress | $12,400 | $28,400 | $340,000 |
| Storage (all tiers) | $34,000 | $55,000 | $420,000 |
| Power (PUE 1.15) | $124,000 | $310,000 | $4,600,000 |
| Data Processing | $20,000 | $35,000 | $200,000 |
| Experimentation/Ablation | $3,500,000 | $5,200,000 | $18,000,000 |
Cost Optimization Strategies
The largest cost-saving opportunity is GPU utilization improvement. Increasing utilization from 65% to 80% on an H200 cluster reduces effective compute cost by 19%. Key levers: optimal parallelism strategy (FSDP for dense models, expert parallelism + FSDP for MoE), efficient data loading (GPUDirect Storage with parallel filesystem), and overlapping communication with computation (using NCCL's multi-threaded all-reduce with computation overlapping). Each percentage point of utilization gain saves approximately $42,700 on a $4.27M GPU compute bill.
Checkpoint frequency optimization is the second-largest opportunity. Reducing checkpoint frequency from every 2 hours to every 4 hours halves checkpoint I/O costs from $5,040 to $2,520 per run, with minimal risk. The risk of a training failure that requires rollback is approximately 2-3% per day for production training runs. At 4-hour checkpoint intervals, the maximum rollback is 4 hours of training, costing approximately $23,800. The optimal checkpoint interval balances recovery cost against storage I/O cost and is typically 3-4 hours for production runs.
