GPU RACK LAYOUT AND CONFIGURATION
GPU rack layout must account for power, cooling, and networking constraints. Standard 42U or 48U racks for GPU clusters: 8x DGX H100 nodes (4 GPUs each) occupy 24U leaving 18U for networking switches, storage, and cable management. Hot aisle containment configuration with cold aisle at 18-22 degrees Celsius and hot aisle at 35-45 degrees Celsius maximizes cooling efficiency.
Weight distribution matters: a fully loaded H100 rack weighs 800-1,200 kg exceeding typical raised floor load ratings of 500-750 kg. Slab floor mounting or weight distribution plates rated for 1,500 kg are required. Power distribution within rack: 2x 60A 480V 3-phase circuits per rack, each feeding 4-8 PSUs. Redundant A/B power feeds prevent single circuit failure causing rack-level outage.
| Component | Per Rack Quantity | Power Draw | Weight | U Space | Cooling Required |
|---|---|---|---|---|---|
| DGX H100 (8-GPU) | 8 nodes | 56 kW | 640 kg | 32U | 50 L/min liquid each |
| InfiniBand switch | 4 switches | 4 kW | 40 kg | 4U | Air (front-back) |
| Management switch | 2 switches | 0.4 kW | 10 kg | 1U | Air |
| Storage node | 2 nodes | 2 kW | 30 kg | 2U | Air |
| PDU | 2 units | 0 kW | 15 kg | 0U | Air |
CABLING BEST PRACTICES
GPU cluster cabling is the most complex physical design element. A 128-node H100 cluster requires: 512 InfiniBand cables (400G each), 256 power cables (C19 to C21), 128 ethernet cables (management), and 128 liquid cooling hoses. Cabling at 8,000+ connections per cluster makes labeling and documentation essential. Color-coded cable management with structured overhead cable trays reduces installation errors by 70 percent.
InfiniBand cabling topology: leaf-spine with 8 leaf switches and 2 spine switches for 128 nodes. Transceiver costs at $400-800 per 400G port add $32,000-$64,000 per cluster. Active Optical Cables cost less than transceivers for distances above 5 meters. Cable minimum bend radius of 30mm for AOC prevents signal degradation.
COOLING INFRASTRUCTURE DESIGN
GPU cooling infrastructure choice depends on rack density. Below 20 kW/rack: standard CRAC air cooling at 150-250 CFM per kW. 20-40 kW/rack: rear-door heat exchangers with water at 15-20 degrees Celsius. Above 40 kW/rack: direct-to-chip liquid cooling with CDU providing 15-25 degrees Celsius water to cold plates. Each 100 kW of GPU load requires approximately 30-40 liters/minute coolant flow.
Cooling redundancy: N+1 for CDUs, 2N for facility chillers. Cooling system PUE contribution: air cooling adds 0.15-0.25 PUE overhead, liquid cooling adds 0.05-0.10. A 1 MW GPU cluster at PUE 1.15 vs 1.30 saves $105,000 annually in power costs. Water consumption for evaporative cooling: 0.5-1.0 gallons per kWh versus 0 for liquid cooling with dry cooler.
SITE SELECTION AND FACILITY REQUIREMENTS
GPU data center site selection evaluates: power availability (10-100 MW from utility), fiber connectivity with sub-10ms latency to major cloud on-ramps, cooling water availability (0.5-2 million gallons/day for evaporative cooling), seismic zone classification, and flood risk (FEMA 100-year flood plain exclusion). Construction timeline: 18-36 months for greenfield, 6-12 months for colocation lease.
Cost per square foot for GPU-optimized data center: colocation $20-$30/square foot/month, retrofit existing facility $500-$800/square foot, greenfield construction $1,000-$1,500/square foot. Total facility cost for 10 MW GPU capacity: $30-$50 million depending on location and cooling technology.
