All essays
MarketMARKET REPORTFEB 2026

CoreWeave GPU Provider Deep Dive: Kubernetes-Native Cloud, NVIDIA Partner, and Pricing Strategy

In-depth analysis of CoreWeave

01

THE COREWEAVE STORY: FROM CRYPTO TO AI INFRASTRUCTURE

CoreWeave began in 2017 as a cryptocurrency mining operation called Atlantic Crypto, accumulating a large GPU fleet before pivoting to cloud rendering and, by 2020, AI workloads. The company rebranded to CoreWeave in 2021 and has since raised over $1.2 billion in debt and equity, including a $2.3 billion debt facility led by Magnetar Capital and Blackstone in 2024. Their infrastructure now spans 14 data center locations across North America and Europe, with over 45,000 NVIDIA H100 and H200 GPUs deployed as of Q2 2026.

The company's key differentiator is its Kubernetes-native architecture. Unlike AWS, GCP, and Azure, which abstract GPU access through proprietary orchestration layers, CoreWeave runs Kubernetes on bare metal with no virtualization overhead. Every GPU node is a Kubernetes worker, providing sub-millisecond pod startup times, direct GPU passthrough via NVIDIA GPU Operator, and native support for Kubernetes device plugins. This architecture delivers 5-15 percent better training throughput versus virtualized alternatives because there is no hypervisor tax on GPU memory access and NVLink interconnects.

02

GPU INVENTORY AND AVAILABILITY

CoreWeave's GPU fleet is concentrated on NVIDIA's highest-end hardware. Their primary offering is the H100 SXM (80 GB HBM3, 3.35 TB/s bandwidth), available in configurations from 1 to 32 GPUs per node via NVLink. They also offer H200 (141 GB HBM3e, 4.8 TB/s) in 8-GPU DGX configurations and began deploying B200 (192 GB HBM3e) in Q1 2026. A100-80G remains available for cost-sensitive workloads, priced approximately 40 percent below H100.

Availability is CoreWeave's competitive edge. During the H100 shortage of 2023-2024, CoreWeave maintained 4-6 week lead times for 64-GPU clusters while hyperscalers quoted 12-20 weeks. As of 2026, H100 lead times are 1-2 weeks, H200 at 2-3 weeks, and B200 at 4-6 weeks. Their inventory advantage stems from preferential allocation as an NVIDIA Elite Partner-one of fewer than 10 partners globally at that tier-which gives them earlier access to new GPU generations and larger allocation volumes.

GPU TypeVRAMMemory BWPricing (On-Demand/hr)Pricing (Reserved 12mo/hr)Availability
H100 SXM80 GB HBM33.35 TB/s$2.75-$3.25$1.85-$2.151-2 weeks
H200 SXM141 GB HBM3e4.8 TB/s$3.50-$4.00$2.40-$2.802-3 weeks
B200 SXM192 GB HBM3e8.0 TB/s$5.00-$6.00$3.60-$4.204-6 weeks
A100-80G SXM80 GB HBM2e2.0 TB/s$1.60-$2.00$1.10-$1.35Available now
L40S48 GB GDDR6864 GB/s$0.85-$1.10$0.55-$0.70Available now
03

KUBERNETES-NATIVE GPU ARCHITECTURE

CoreWeave's infrastructure is built on a custom Kubernetes distribution running on bare metal, not on top of ESXi or KVM. Each GPU node runs Kubelet with the NVIDIA GPU Operator managing driver installation, device plugin registration, and MIG (Multi-Instance GPU) configuration. The control plane uses a multi-cluster management layer that maps Kubernetes namespaces to tenant accounts, resource quotas, and network policies.

What distinguishes this approach is the networking layer. CoreWeave deploys a custom CNI plugin that provides direct node-to-node GPU communication over the cluster's InfiniBand or RoCE v2 fabric, bypassing the Kubernetes overlay network entirely for training traffic. This means that a pod running distributed training over PyTorch DDP or FSDP experiences the same inter-node latency as a bare metal deployment: under 2 microseconds for NVLink within a node and under 10 microseconds for InfiniBand across nodes. The result is 97-99 percent of bare metal training throughput, a figure confirmed by internal CoreWeave benchmarks and by third-party tests from companies like Together AI and Anyscale.

04

PRICING MODEL AND RESERVATION STRUCTURE

CoreWeave offers three pricing tiers: on-demand (hourly, no commitment), reserved (1-month to 36-month commitments), and spot instances (preemptible, up to 70 percent discount). On-demand H100 pricing ranges from $2.75-3.25 per GPU-hour, comparable to hyperscaler on-demand rates. The value emerges at the reserved level: a 12-month H100 commitment drops to $2.00 per GPU-hour, while 36-month commitments reach $1.50 per GPU-hour. Spot instances fluctuate between $0.60-1.20 per GPU-hour with 5-15 percent interruption rates.

CoreWeave differentiates with dynamic resource bundling: a single reservation can include a mix of GPU types, and customers can adjust the ratio monthly within their committed total. For example, a $100,000/month reserved customer can run 40 H100s in January and 30 H100s plus 8 B200s in February without renegotiating. This flexibility is unique among GPU providers and addresses the reality that AI companies' GPU requirements shift as they move from experimentation to training to production inference.

Commitment LevelH100/hrH200/hrB200/hrA100/hrSavings vs On-Demand
On-Demand$3.00$3.75$5.50$1.800%
3-Month Reserved$2.40$3.00$4.40$1.4520%
12-Month Reserved$2.00$2.60$3.90$1.2533%
36-Month Reserved$1.50$2.00$3.00$0.9550%
Spot (Preemptible)$0.90$1.10TBD$0.5070%
05

DATA CENTER FOOTPRINT AND FUTURE EXPANSION

CoreWeave operates 14 data centers across the US and Europe as of Q2 2026. US locations include Las Vegas (NV), Chicago (IL), Dallas (TX), Ashburn (VA), Charlotte (NC), Atlanta (GA), and Seattle (WA). European locations span London (UK), Frankfurt (DE), Amsterdam (NL), Stockholm (SE), and Paris (FR), with Dublin (IE) and Milan (IT) under construction. This geographic distribution allows sub-20ms latency for most US and Western European users, though APAC remains a notable gap.

The company announced a $1.5 billion expansion plan for 2026-2027, adding 8 new data centers, including two in APAC (Tokyo and Singapore) and one in South America (Sao Paulo). The expansion is primarily B200-focused, with 35,000 B200 GPUs on order. CoreWeave has also invested in renewable energy: its data center portfolio sources 60 percent of power from wind and solar, with a target of 100 percent by 2028.

RegionData CentersPrimary GPUNetworkPower (Total MW)
US WestLas Vegas, SeattleH100, B200InfiniBand NDR85 MW
US CentralChicago, DallasH100, H200InfiniBand NDR70 MW
US EastAshburn, Charlotte, AtlantaH100, H200, B200InfiniBand NDR120 MW
EuropeLondon, Frankfurt, Amsterdam, Stockholm, ParisH100, L40SInfiniBand HDR95 MW
APAC (2027)Tokyo, SingaporeB200InfiniBand NDR50 MW
06

COMPETITIVE POSITION VS HYPERSCALERS AND PEER PROVIDERS

CoreWeave occupies a middle ground between hyperscalers (AWS, GCP, Azure) and smaller GPU providers (Lambda, RunPod, Vast.ai). Their pricing at the 12-month reserved level is 20-35 percent below hyperscaler on-demand rates and roughly on par with hyperscaler 1-year commitments. The Kubernetes-native architecture provides a meaningful advantage for teams that already run on Kubernetes: they can migrate workloads between CoreWeave and their own on-prem clusters with minimal configuration changes.

However, CoreWeave lacks the managed services that hyperscalers offer: no managed databases, no managed inference endpoints (beyond basic Kubernetes), and no integrated AI platform like Vertex AI or SageMaker. For well-staffed AI companies with infrastructure engineers, this is a feature. For smaller teams looking for a turnkey solution, CoreWeave adds operational overhead. The trade-off is clear: CoreWeave delivers 20-35 percent lower GPU cost for teams that can manage their own Kubernetes stack.

Filed under
CoreWeaveKubernetes GPUNVIDIA PartnerH100 CloudB200 GPUGPU Cloud PricingCloud-Native GPU