All essays
MarketMARKET REPORTFEB 2026

True Cost of LLM Inference in 2025: Cost-Per-Token on H100, H200, and B200

Inference is 70%+ of production GPU spend. Definitive 2025 guide to self-hosted infrastructure cost per token.

01

COST METHODOLOGY

True inference cost includes GPU rental, KV cache memory overhead, networking for multi-GPU deployments, power, and operational overhead. We calculate cost-per-million-tokens (CPMT) using sustained throughput benchmarks at 50-80% utilization over a 30-day period. API pricing from providers is compared against self-hosted infrastructure to determine the breakeven volume where self-hosting becomes economical.

02

H100 COST-PER-TOKEN

H100 serving Llama 3 70B at FP8 achieves 3,800 tok/s with 8 concurrent requests. At $2.50/hr spot: $0.18/M tokens output. With 70% utilization over 30 days: $0.13/M tokens effective. H100's 80 GB limits batch size to 8-10 requests for 70B models at 4K context, capping throughput at 4,200 tok/s. Cost advantage degrades at longer context lengths where KV cache reduces effective batch size.

ModelGPUThroughputCost/hrCPMT
Llama 3 8B1x L42,100 tok/s$0.50$0.07
Llama 3 70B1x H1003,800 tok/s$2.50$0.18
Llama 3 70B1x H2005,200 tok/s$3.20$0.17
Llama 3 405B8x H1006,500 tok/s$20.00$0.85
Llama 3 70B1x B2007,500 tok/s$4.50$0.17
03

H200 COST-PER-TOKEN

H200's 141 GB enables 2x batch size versus H100 for 70B models at 4K context: 10,400 tok/s versus 3,800 tok/s. At $3.20/hr reserved: $0.085/M tokens-50% improvement over H100. For 32K context workloads, H200 advantage grows to 2.5-3x as KV cache fits single GPU. H200 is the most cost-effective Hopper GPU for long-context inference.

04

B200 COST-PER-TOKEN

B200 at FP8 achieves 7,500 tok/s for 70B models, 2x H100 throughput. At $4.50/hr: $0.17/M tokens-similar to H100 cost-per-token. B200's advantage emerges with NVFP4: 15,000 tok/s at $0.08/M tokens, 2.1x improvement over H100 FP8. For large batches and long context, B200's 192 GB delivers 3-4x effective throughput versus H100.

05

SCALE ECONOMICS

At 100M tokens/day: H100 costs $19,500/month, B200 costs $25,500/month, H200 costs $15,500/month. At 1B tokens/day: H100 cluster (10 GPUs) costs $195,000/month, B200 cluster (5 GPUs with FP4) costs $120,000/month. Breakeven between API and self-host: 50-100M tokens/month for 70B models on H100, 20-50M tokens/month for 8B models on L4.

06

PROVIDER COMPARISON

Lambda offers H100 at $2.50/hr spot with 50% utilization guarantee. CoreWeave provides H200 reserved at $3.00-3.50/hr with NVLink. RunPod B200 at $3.80-4.50/hr spot is best Blackwell value. Vast.ai H100 at $1.03/hr spot is the cheapest but with variable availability. For production inference, reserved contracts with 90%+ utilization guarantee are worth the 15-25% premium over spot.

Filed under
LLM CostInference 2025Cost Per TokenH100H200B200Inference Pricing