Why the Sticker Price Is Only 40% of What You Pay
When a GPU provider quotes $3.36/GPU/hr for a B200 or $2.02/GPU/hr for an H200, that number includes only the hardware amortization, facility base load, and a margin. The remaining costs land on your P&L as separate line items: networking, storage, power, cooling, facility space, staffing, software licensing, and downtime. In a fully loaded TCO calculation, these hidden costs add 50-150% to the raw GPU rental rate depending on deployment model.
The split varies dramatically by deployment model. In hyperscaler cloud (AWS, Azure, GCP), most infrastructure costs are absorbed into the per-hour rate, but you pay via egress fees ($0.05-0.12/GB), storage costs ($0.02-0.08/GB/month for object storage), and software licensing markups. In bare-metal colocation, you pay the GPU lease plus facility power and space separately. In on-premise deployment, you own every cost category from hardware depreciation to cooling tower maintenance. The same GPU workload can have 3x different total costs across these models.
This post breaks down each hidden cost category with mid-2026 pricing data, then builds a full TCO model for three common deployment scenarios: 8x H200 node at a colocation provider, equivalent node in AWS with reserved capacity, and a self-owned 8x H200 server in a company data center. We include real ranges so you can adjust for your geography and provider relationships.
Networking: The $500-$5,000 Per Port Line Item
Networking costs are the most frequently omitted item in GPU cluster TCO models. Every GPU node needs at least one high-speed network connection for data transfer and typically a second for storage access. For a multi-node cluster, you need top-of-rack switches, inter-rack cabling, and sometimes spine switches depending on cluster size. The per-port cost ranges from roughly $500 for a 100GbE Ethernet port to $5,000+ for NDR400 InfiniBand with optical transceivers.
For a 64-GPU cluster (8 DGX B200 nodes, each with 8 GPUs), the networking cost breaks down as follows. Eight nodes each need 4x InfiniBand NDR400 ports for GPU-to-GPU communication (32 ports at $4,200 each = $134,400), plus 2x 100GbE ports per node for storage access and management (16 ports at $800 each = $12,800). Add one Quantum-3 NDR400 InfiniBand switch with 32 ports at $120,000 and two 100GbE switches at $15,000 each. Total networking hardware: approximately $290,000, or $4,530 per GPU.
For teams renting GPU time from a bare-metal provider, networking cost shows up indirectly in the per-GPU price. Providers running Spectrum-X Ethernet can pass 25% lower network costs than those running InfiniBand. Our GPU network cost analysis covers per-port pricing across all three fabric types from $500 to $5,000+.
| Networking Component | Per-Unit Cost | Qty for 64 GPUs | Extended Cost |
|---|---|---|---|
| InfiniBand NDR400 port (CX-8) | $4,200 | 32 | $134,400 |
| 100GbE port (ConnectX-7) | $800 | 16 | $12,800 |
| Quantum-3 NDR400 managed switch | $120,000 | 1 | $120,000 |
| 100GbE top-of-rack switch | $15,000 | 2 | $30,000 |
| Optical transceivers + cabling | $180 | 48 | $8,640 |
| Total networking capex per GPU | $4,530 |
Power and Cooling: The Recurring Cost That Doubles Your Bill
Power is the largest hidden recurring cost in GPU operations. An H200 SXM5 at full load draws 700W. Add CPU, memory, networking, and fans in a standard server chassis, and a fully loaded 8xH200 node draws approximately 7.5-8.0 kW at the wall. At the US commercial average of $0.12/kWh, that is $0.96/node/hr or $0.12/GPU/hr in electricity alone. Over a 3-year deployment at 80% utilization, power cost per GPU is approximately $2,520 - nearly equal to the GPU's residual hardware value.
Cooling adds another 30-60% on top of IT power consumption depending on cooling technology. Air-cooled data centers at 15-20 kW per rack (typical for H200 clusters) have a PUE (Power Usage Effectiveness) of 1.3-1.6, meaning every watt of IT power requires 0.3-0.6 watts of cooling and facility overhead. For the 8xH200 node at 8kW IT power and PUE 1.5, total facility power is 12kW. At $0.12/kWh, total power + cooling is $12.64/day or $0.066/GPU/hr in overhead beyond the baseline hardware rental. Liquid-cooled deployments for B200 or B300 clusters at 40-140 kW per rack can achieve PUE of 1.05-1.15, reducing the cooling overhead by 60-80% but adding capex for CDUs and fluid distribution.
Geography dramatically changes the power cost. Norway at $0.04-0.06/kWh or Quebec at $0.05-0.07/kWh produce power costs one-third to one-half of California at $0.18-0.25/kWh or Germany at $0.20-0.35/kWh. For a 64-GPU H200 cluster at full load, the difference between Norway and California power is roughly $50,000-70,000/year in operating cost. See our GPU data center location strategy for geographic power price analysis.
| GPU Model | TDP | Node Power (8-GPU) | Cooling PUE | Total Power @$0.12/kWh |
|---|---|---|---|---|
| H100 SXM5 | 700 W | 7.0 kW | 1.5 | $1.04/node/hr |
| H200 SXM5 | 700 W | 7.5 kW | 1.5 | $1.12/node/hr |
| B200 SXM6 | 1,000 W | 10.5 kW | 1.2 | $1.26/node/hr |
| B300 SXM7 | 1,400 W | 14.0 kW | 1.1 | $1.54/node/hr |
| MI300X | 750 W | 8.0 kW | 1.5 | $1.20/node/hr |
Storage: The Parallel Filesystem Tax
AI workloads consume storage at two tiers: high-performance parallel filesystem for training data during compute, and lower-cost object or file storage for datasets at rest. A 64-GPU H200 training cluster typically needs 200-500 TB of NVMe-backed parallel storage (Lustre, WekaFS, VAST Data, or DAOS) delivering 40-80 GB/s read throughput to avoid GPU data starvation. At mid-2026 pricing, parallel filesystem as-a-service runs $0.08-0.18/GB/month provisioned, or $1,600-3,600/month for 20 TB of high-performance capacity.
The storage cost trap is provisioning for peak throughput. Training data loaders are bursty - you need high throughput during data loading but the storage sits mostly idle during compute. Most providers provision storage at peak capacity. On a 64-GPU H200 cluster running 24/7, storage adds roughly $4.50-8.50/GPU/hr in effectively hidden cost when you amortize the parallel filesystem cost over total provisioning. Our parallel filesystem tax analysis covers the full breakdown.
Object storage for cold datasets adds another layer. 1 PB of S3-compatible object storage at $0.006-0.015/GB/month costs $6,000-15,000/month. Data transfer to the parallel filesystem tier adds egress charges within and across clouds. Egress from object storage to compute within the same region is usually free; cross-region egress can add $0.01-0.09/GB transferred.
Staffing: The Cost That Never Appears on an Invoice
Staffing is consistently the most underestimated GPU TCO component because it is not itemized in any provider quote. A production GPU cluster requires at minimum: one infrastructure engineer per 32-64 GPUs for cluster management, Kubernetes, networking, and storage operations; one ML engineer per 16-32 GPUs for framework optimization, model deployment, and performance tuning; and fractional SRE and security coverage. At mid-2026 US salaries ($160K-220K fully loaded for senior infra roles), staffing adds $50-70/GPU/month in a colocation or on-premise model.
In cloud or bare-metal rental, staffing requirements do not disappear - they shift. The cloud provider handles hardware maintenance, power, and cooling, but your team still handles OS patching, container orchestration, model serving setup, monitoring, and capacity management. Cloud-managed services (SageMaker, Vertex AI, Azure ML) reduce staffing needs but add 20-40% margin on underlying compute costs. Our analysis of teams operating 64-256 GPU clusters across models shows that staffing represents 18-32% of total TCO in colocation/on-prem and 8-15% in fully managed cloud.
The staffing cost also varies by GPU platform. ROCm-based clusters (MI300X) require more engineer time for kernel debugging and performance optimization, adding roughly 25-40% more engineering hours compared to CUDA-based clusters at the same GPU count. This is a soft cost that rarely makes it into procurement decks but shows up in the velocity of ML research output.
| Role | Ratio | Monthly Cost (FTE) | Per 64-GPU Cluster |
|---|---|---|---|
| Infrastructure Engineer | 1 per 32-64 GPUs | $15,000-18,000 | $15,000-30,000 |
| ML Engineer / Performance | 1 per 16-32 GPUs | $17,000-20,000 | $34,000-80,000 |
| SRE (fractional) | 0.5 per 64 GPUs | $8,500-10,000 | $4,250-5,000 |
| Storage Admin (fractional) | 0.25 per 64 GPUs | $15,000-17,000 | $3,750-4,250 |
| Security (fractional) | 0.1 per 64 GPUs | $16,000-20,000 | $1,600-2,000 |
| Total staffing per month | $58,600-121,250 |
Software Licensing and Support That Quietly Accumulates
Software costs are the stealthiest TCO line item because they arrive as procurement requests months after the hardware is installed. For CUDA-based clusters, the largest cost is typically NVIDIA AI Enterprise licensing, which covers TensorRT-LLM, NVIDIA Dynamo, and NVIDIA NIM. At $4,500-9,000/GPU/year for production deployments, a 64-GPU cluster running NVIDIA AI Enterprise adds $24,000-48,000/month in licensing. Many teams run without this license (using open-source vLLM or SGLang), but lose access to TensorRT-LLM's highest-throughput inference and NVIDIA Dynamo's disaggregated serving.
Other software costs include: Lustre or WekaFS licensing ($15,000-25,000/year per 100 TB), monitoring tools (Grafana Cloud at $2,000-10,000/month for enterprise GPU monitoring), container registry and CI/CD infrastructure ($1,000-5,000/month), and ML platform tools (MLflow, Weights & Biases, or Neptune at $1,500-15,000/month depending on seat count and artifact storage). These costs add up to $8,000-20,000/month for a 64-GPU cluster, or roughly $0.17-0.42/GPU/hr.
Open-source alternatives reduce this line item significantly. Running vLLM instead of TensorRT-LLM, MinIO instead of licensed parallel filesystem, and self-hosted Prometheus/Grafana instead of Grafana Cloud can cut software costs by 60-80%. The trade-off is engineering time to set up, maintain, and optimize these stacks, which brings back to the staffing cost discussion.
| Software Category | Enterprise Option | Annual Cost (64 GPU) | Open Source Alternative |
|---|---|---|---|
| Inference framework | NVIDIA AI Enterprise | $288,000-576,000 | vLLM / SGLang (free) |
| Parallel filesystem | WekaFS / Lustre license | $15,000-25,000 | MinIO (partial) |
| Observability | Grafana Cloud / Datadog | $24,000-120,000 | Self-hosted Prometheus + Grafana |
| ML platform | W&B / Neptune enterprise | $18,000-60,000 | MLflow (open source) |
| Container registry | Docker Hub / ECR | $12,000-48,000 | Harbor (self-hosted) |
Full TCO Comparison: Cloud vs Colocation vs On-Premise
Below is the fully loaded monthly TCO for an 8x H200 node (64 GPUs total across 8 nodes) in three deployment models. Cloud: AWS p5.48xlarge instances with 3-year reserved pricing at $18.24/hr per instance (8 instances). Colocation: H200 hosted at a US neocloud provider, $2.02/GPU/hr plus power at $0.12/kWh metered. On-premise: purchased H200 servers amortized over 4 years ($220K per node, $1.57/GPU/hr depreciation), plus facility power, cooling, and staffing at standard colocation rates.
The cloud model appears most expensive on raw GPU cost but requires the least staffing and includes all facility costs. Colocation is the middle option: lower GPU hardware cost than on-premise, but power and staffing are separate. On-premise has the lowest marginal GPU cost after the initial capital outlay, but requires the most staffing, facility space, and upfront capital. The break-even between cloud and on-premise occurs at roughly 65-75% GPU utilization over 3 years, consistent with our buying vs renting TCO analysis.
The key insight: colocation and on-premise look cheaper on a per-GPU-hour basis, but only if you have the team and operational maturity to handle the non-GPU costs. Many AI teams that start in colocation end up spending 30-50% more than budgeted because they underestimated networking, storage, and staffing. ClusterBid transparently factors these costs into quotes and can help you compare total cost across deployment models. For current pricing across providers, see the inventory page.
| Cost Category | Cloud (AWS, 3yr) | Colocation (8-node) | On-Premise (owned) |
|---|---|---|---|
| GPU compute (64 GPUs) | $43,776 | $92,928 | $72,192 |
| Power + cooling (included) | Included | $18,432 | $18,432 |
| Networking amortized | Included | $4,028 | $4,028 |
| Storage (parallel FS 50TB) | $4,800 | $6,000 | $3,000 |
| Staffing (infra + ML) | $28,000 | $48,000 | $52,000 |
| Software licensing | $12,000 | $8,000 | $24,000 |
| Facility space | Included | $3,200 | $1,600 |
| Total monthly TCO | $88,576 | $180,588 | $175,252 |
| Effective per GPU/hr | $1.92 | $3.92 | $3.80 |
