All essays
TechnicalDEEP DIVEFEB 2026

Liquid Cooling and GPU Data Centers: What Blackwell Buyers Must Verify Before Signing a Contract

B200 requires liquid cooling at 1200W TDP. Most older data centers are air-cooled to 700W max per server slot. Detailed guide on what to verify before signing a colocation or lease contract.

01

BLACKWELL COOLING REQUIREMENTS

NVIDIA's B200 GPU carries a 1200W TDP, more than double the H100's 700W. A standard DGX B200 node with 8 GPUs draws 10-12 kW per rack unit, far exceeding the 15-25 kW per rack typical of air-cooled data centers designed for CPU workloads. The thermal density requires direct-to-chip liquid cooling with water temperatures of 35-45C inlet and coolant flow rates of 2-4 liters per minute per GPU.

B300 pushes further to 1500W TDP per GPU, requiring reinforced liquid cooling infrastructure. The transition from air to liquid cooling is not optional for Blackwell-class GPUs-NVIDIA specifies liquid cooling as mandatory for B200 deployments, with air cooling unsupported for sustained operation above 500W per GPU.

02

LIQUID COOLING TECHNOLOGY OPTIONS

Three liquid cooling approaches serve Blackwell clusters. Direct-to-chip (cold plate) cooling circulates coolant through plates mounted directly on GPU packages, removing 80-90% of heat at the source with the remaining 10-20% handled by facility air. This is the most common approach for Blackwell deployments, with solutions from CoolIT, Boyd, and Asia Vital Components.

Immersion cooling submerges entire servers in dielectric fluid, eliminating air cooling entirely. This supports higher power densities (100+ kW per rack) but complicates GPU maintenance and upgrades. Single-phase immersion (using fluorocarbon fluids) is the safer choice for GPU clusters, while two-phase immersion (boiling fluid) offers higher efficiency but adds complexity. Liquid-cooled rear-door heat exchangers are a third, less effective option suitable only for clusters below 50 kW per rack.

03

FACILITY VERIFICATION CHECKLIST

Before signing any colocation or lease contract, verify: facility water supply temperature (must be 45C or below for B200), available water pressure (2-4 bar minimum at rack location), condensation control systems for supply lines, and floor load rating (minimum 2,500 lbs per rack for liquid-cooled GPU clusters versus 1,500 lbs for standard). Most Tier 3 data centers built before 2024 cannot meet these requirements without retrofit.

Additional verification items include: redundant cooling loops (N+1 minimum, 2N preferred), dielectric fluid leak detection with automatic shutoff, humidity control systems capable of maintaining 40-60% RH at elevated temperatures, and fire suppression rated for liquid-cooled environments (gas-based suppression instead of water sprinklers).

04

COST IMPLICATIONS OF LIQUID COOLING

Liquid cooling adds $0.05-$0.12 per GPU-hour to infrastructure costs versus air cooling. For a 72-GPU B200 cluster running 24/7, this adds $315-$756 per month in cooling costs. The capital expenditure for liquid cooling retrofit is $3,000-$8,000 per rack for direct-to-chip systems including coolant distribution units, manifolds, and installation.

The TCO comparison favors liquid cooling despite higher upfront costs when GPU density is factored. Liquid cooling enables 40-80 kW per rack versus 15-25 kW for air, reducing data center floor space requirements by 50-70%. For clusters exceeding 100 GPUs, the floor space savings offset the liquid cooling premium within 12-18 months.

05

CONTRACT PROTECTION CLAUSES

GPU colocation contracts for Blackwell hardware require specific protections. Include: guaranteed water supply temperature (with SLA credits for supply water exceeding 45C), coolant quality specifications with testing rights, maximum power interruption duration (sub-5-minute UPS coverage mandatory), and liquid cooling maintenance response times (4-hour on-site for cooling loop failures).

Force majeure exclusions should specifically address cooling system failures and water supply interruptions-standard force majeure clauses often exclude these as the facility's responsibility. Include rights to deploy alternative cooling methods at provider expense if the installed system fails to maintain GPU operating temperature below 85C junction.

06

PROVIDER EVALUATION

Evaluate colocation providers on existing liquid cooling deployment experience. Providers with operating Blackwell or H100 liquid-cooled clusters (CoreSite, Equinix, Digital Realty with retrofits) have demonstrated facilities competence. Avoid providers for whom your deployment would be their first liquid-cooled installation-the learning curve risk is substantial.

Request reference calls with three existing liquid-cooled GPU customers of the provider. Verify: average coolant supply temperature by month (thermal excursions are most common in summer months), actual PUE achieved (should be 1.15-1.25 for liquid-cooled deployments versus 1.3-1.5 for air-cooled), and mean time between cooling-related incidents. A provider unable to provide these references should be avoided for Blackwell deployment.

Filed under
Liquid CoolingBlackwell CoolingB200 CoolingGPU Data CenterCooling TCOData CenterInfrastructure