Tier 1: Enterprise-Grade Providers
CoreWeave and Lambda occupy Tier 1 for good reason. CoreWeave operates over 50 data center locations with 45,000+ H100 and B200 GPUs under management. Their Kubernetes-native infrastructure, with automated node recovery, persistent storage classes, and direct NVLink-connected GPU pods, sets the standard for production AI infrastructure. H200 spot pricing averages $2.85/hr with 99.99% uptime SLA on reserved contracts.
Lambda Cloud offers a more curated experience with fewer SKUs but excellent execution. Their GPU clusters come pre-configured with optimized networking (Spectrum-X Ethernet at 400 Gbps) and pre-installed CUDA 13.1, PyTorch, and NCCL. Lambda's support team responds within 15 minutes on business-critical tickets, a stark contrast to most providers. B200 pricing at $4.20/hr spot is competitive with larger providers.
Tier 2: Scale-Up Neoclouds
RunPod, Vast.ai, and Paperspace form Tier 2: providers that have scaled beyond hobbyist infrastructure but still show operational inconsistencies. RunPod offers the most granular GPU selection (from RTX 4090s to H100s) with serverless GPU endpoints that auto-scale based on request volume. Their spot pricing on H100 is among the lowest at $1.89/hr, but provisioning delays of 10-30 minutes during peak hours are common.
Vast.ai operates the largest peer-to-peer GPU marketplace, aggregating GPUs from individual hosts and small data centers. This gives Vast the widest geographic distribution and lowest prices (H100s at $1.65/hr), but reliability varies dramatically by host. GPU failures, network disconnects, and inconsistent performance are frequent enough that Vast is unsuitable for production training workloads without checkpoint redundancy.
Tier 3: Specialized and Regional Providers
Crusoe Cloud, TensorDock, FluidStack, and DataCrunch form Tier 3: providers with specific geographic or technical advantages but limited scale. Crusoe runs on stranded methane gas in data centers in Colorado and Oklahoma, offering carbon-negative claims that appeal to ESG-conscious teams. Their H100 pricing at $2.40/hr is competitive, but their total capacity under 5,000 GPUs limits availability for large jobs.
TensorDock specializes in European GPU hosting with data centers in Iceland, Norway, and the Netherlands. H100s run on 100% renewable hydroelectric power at $2.20/hr, 15-25% below US prices. DataCrunch offers bare-metal H200 nodes in Northern Virginia and Frankfurt with 24-hour provisioning and direct-dial support, making them a strong choice for teams that need rapid deployment.
| Provider | H200 Spot/hr | B200 Spot/hr | Uptime SLA | Provisioning | Support Tier |
|---|---|---|---|---|---|
| CoreWeave | $2.85 | $4.95 | 99.99% | <5 min | 24/7 engineer |
| Lambda | $3.10 | $4.20 | 99.95% | <10 min | 15-min response |
| RunPod | $1.89 | $3.50 | 99.9% | 10-30 min | Ticket only |
| Vast.ai | $1.65 | N/A | 99.5% | 2-15 min | Community/chat |
| Paperspace | $2.50 | $5.10 | 99.95% | <5 min | 24/7 chat |
| Crusoe | $2.40 | N/A | 99.9% | <30 min | Business hours |
| TensorDock | $2.20 | $3.80 | 99.8% | <15 min | Email/ticket |
| DataCrunch | $2.60 | $4.50 | 99.95% | <24 hr | Direct dial |
Tier 4: Hyperscalers
AWS, Azure, and GCP occupy their own tier. Their GPU offerings are generally 30-60% more expensive than neoclouds on on-demand pricing, but they offer unmatched ecosystem integration, regulatory compliance certifications, and geographic reach. AWS P5 instances with H100s run at $14.35/hr on-demand versus $4.20/hr for equivalent capacity on CoreWeave.
The hyperscaler advantage is distribution and procurement simplicity. AWS offers H200 capacity in 14 regions, Azure in 12, and GCP in 10. For teams that need multi-region deployment with HIPAA, SOC 2, or FedRAMP compliance, the hyperscaler premium is often unavoidable. Reserved instances (1-3 year terms) narrow the gap to roughly 15-25% above neocloud pricing.
Evaluation Criteria and Methodology
Our tier list is based on four weighted criteria: effective price (pricing minus hidden fees for egress, storage, and API calls), infrastructure reliability (GPU failure rate, network stability, provisioning success rate), support quality (response time, resolution rate, escalation path), and deployment velocity (time from signup to first training job).
Hidden fees substantially alter the effective price. CoreWeave includes 10 TB of free egress per month per node; Lambda includes 5 TB. RunPod charges $0.05/GB for egress beyond 1 TB. Vast.ai charges $0.03/GB for storage above the included 10 GB. These fees can add 10-30% to the effective cost for data-heavy workloads.
Recommendations by Use Case
For production training at scale (64+ GPUs, 24/7 operation): CoreWeave or Lambda. Both offer the infrastructure maturity, support responsiveness, and reliability guarantees that sustained training requires. CoreWeave's Kubernetes-native tooling gives it a slight edge for teams with existing K8s expertise.
For inference serving with variable demand: RunPod's serverless GPU endpoints provide auto-scaling with millisecond cold starts. Their per-second billing eliminates waste during idle periods. For teams with steady-state inference demand, Lambda's reserved instances at $3.40/hr for H200 offer better TCO.
For GPU prototyping and experimentation: Vast.ai or RunPod spot instances minimize costs for non-critical workloads. Vast's $1.65/hr H100s are ideal for model evaluation, data preprocessing, and short fine-tuning runs where checkpoint recovery handles the occasional host failure.
