All essays
TechnicalDEEP DIVEFEB 2026

Physical Infrastructure for GPU Clusters: Rack Layout, Power, Cooling

Physical infrastructure design for GPU data centers covering rack layout optimization, power distribution architecture, cooling system selection, and cabling best practices.

01

GPU RACK LAYOUT AND CONFIGURATION

GPU rack layout must account for power, cooling, and networking constraints. Standard 42U or 48U racks for GPU clusters: 8x DGX H100 nodes (4 GPUs each) occupy 24U leaving 18U for networking switches, storage, and cable management. Hot aisle containment configuration with cold aisle at 18-22 degrees Celsius and hot aisle at 35-45 degrees Celsius maximizes cooling efficiency.

Weight distribution matters: a fully loaded H100 rack weighs 800-1,200 kg exceeding typical raised floor load ratings of 500-750 kg. Slab floor mounting or weight distribution plates rated for 1,500 kg are required. Power distribution within rack: 2x 60A 480V 3-phase circuits per rack, each feeding 4-8 PSUs. Redundant A/B power feeds prevent single circuit failure causing rack-level outage.

ComponentPer Rack QuantityPower DrawWeightU SpaceCooling Required
DGX H100 (8-GPU)8 nodes56 kW640 kg32U50 L/min liquid each
InfiniBand switch4 switches4 kW40 kg4UAir (front-back)
Management switch2 switches0.4 kW10 kg1UAir
Storage node2 nodes2 kW30 kg2UAir
PDU2 units0 kW15 kg0UAir
02

CABLING BEST PRACTICES

GPU cluster cabling is the most complex physical design element. A 128-node H100 cluster requires: 512 InfiniBand cables (400G each), 256 power cables (C19 to C21), 128 ethernet cables (management), and 128 liquid cooling hoses. Cabling at 8,000+ connections per cluster makes labeling and documentation essential. Color-coded cable management with structured overhead cable trays reduces installation errors by 70 percent.

InfiniBand cabling topology: leaf-spine with 8 leaf switches and 2 spine switches for 128 nodes. Transceiver costs at $400-800 per 400G port add $32,000-$64,000 per cluster. Active Optical Cables cost less than transceivers for distances above 5 meters. Cable minimum bend radius of 30mm for AOC prevents signal degradation.

03

COOLING INFRASTRUCTURE DESIGN

GPU cooling infrastructure choice depends on rack density. Below 20 kW/rack: standard CRAC air cooling at 150-250 CFM per kW. 20-40 kW/rack: rear-door heat exchangers with water at 15-20 degrees Celsius. Above 40 kW/rack: direct-to-chip liquid cooling with CDU providing 15-25 degrees Celsius water to cold plates. Each 100 kW of GPU load requires approximately 30-40 liters/minute coolant flow.

Cooling redundancy: N+1 for CDUs, 2N for facility chillers. Cooling system PUE contribution: air cooling adds 0.15-0.25 PUE overhead, liquid cooling adds 0.05-0.10. A 1 MW GPU cluster at PUE 1.15 vs 1.30 saves $105,000 annually in power costs. Water consumption for evaporative cooling: 0.5-1.0 gallons per kWh versus 0 for liquid cooling with dry cooler.

04

SITE SELECTION AND FACILITY REQUIREMENTS

GPU data center site selection evaluates: power availability (10-100 MW from utility), fiber connectivity with sub-10ms latency to major cloud on-ramps, cooling water availability (0.5-2 million gallons/day for evaporative cooling), seismic zone classification, and flood risk (FEMA 100-year flood plain exclusion). Construction timeline: 18-36 months for greenfield, 6-12 months for colocation lease.

Cost per square foot for GPU-optimized data center: colocation $20-$30/square foot/month, retrofit existing facility $500-$800/square foot, greenfield construction $1,000-$1,500/square foot. Total facility cost for 10 MW GPU capacity: $30-$50 million depending on location and cooling technology.

Filed under
Data CenterPhysical InfrastructureRack LayoutPower DistributionCablingCooling DesignGPU Facility