COST METHODOLOGY
True inference cost includes GPU rental, KV cache memory overhead, networking for multi-GPU deployments, power, and operational overhead. We calculate cost-per-million-tokens (CPMT) using sustained throughput benchmarks at 50-80% utilization over a 30-day period. API pricing from providers is compared against self-hosted infrastructure to determine the breakeven volume where self-hosting becomes economical.
H100 COST-PER-TOKEN
H100 serving Llama 3 70B at FP8 achieves 3,800 tok/s with 8 concurrent requests. At $2.50/hr spot: $0.18/M tokens output. With 70% utilization over 30 days: $0.13/M tokens effective. H100's 80 GB limits batch size to 8-10 requests for 70B models at 4K context, capping throughput at 4,200 tok/s. Cost advantage degrades at longer context lengths where KV cache reduces effective batch size.
| Model | GPU | Throughput | Cost/hr | CPMT |
|---|---|---|---|---|
| Llama 3 8B | 1x L4 | 2,100 tok/s | $0.50 | $0.07 |
| Llama 3 70B | 1x H100 | 3,800 tok/s | $2.50 | $0.18 |
| Llama 3 70B | 1x H200 | 5,200 tok/s | $3.20 | $0.17 |
| Llama 3 405B | 8x H100 | 6,500 tok/s | $20.00 | $0.85 |
| Llama 3 70B | 1x B200 | 7,500 tok/s | $4.50 | $0.17 |
H200 COST-PER-TOKEN
H200's 141 GB enables 2x batch size versus H100 for 70B models at 4K context: 10,400 tok/s versus 3,800 tok/s. At $3.20/hr reserved: $0.085/M tokens-50% improvement over H100. For 32K context workloads, H200 advantage grows to 2.5-3x as KV cache fits single GPU. H200 is the most cost-effective Hopper GPU for long-context inference.
B200 COST-PER-TOKEN
B200 at FP8 achieves 7,500 tok/s for 70B models, 2x H100 throughput. At $4.50/hr: $0.17/M tokens-similar to H100 cost-per-token. B200's advantage emerges with NVFP4: 15,000 tok/s at $0.08/M tokens, 2.1x improvement over H100 FP8. For large batches and long context, B200's 192 GB delivers 3-4x effective throughput versus H100.
SCALE ECONOMICS
At 100M tokens/day: H100 costs $19,500/month, B200 costs $25,500/month, H200 costs $15,500/month. At 1B tokens/day: H100 cluster (10 GPUs) costs $195,000/month, B200 cluster (5 GPUs with FP4) costs $120,000/month. Breakeven between API and self-host: 50-100M tokens/month for 70B models on H100, 20-50M tokens/month for 8B models on L4.
PROVIDER COMPARISON
Lambda offers H100 at $2.50/hr spot with 50% utilization guarantee. CoreWeave provides H200 reserved at $3.00-3.50/hr with NVLink. RunPod B200 at $3.80-4.50/hr spot is best Blackwell value. Vast.ai H100 at $1.03/hr spot is the cheapest but with variable availability. For production inference, reserved contracts with 90%+ utilization guarantee are worth the 15-25% premium over spot.
