All essays
GuideGUIDEFEB 2026

GPU Cluster Networking Setup: InfiniBand Partitioning and Fabric Management for AI Workloads

Design and manage GPU cluster InfiniBand fabric. Subnet manager configuration, PKey partitioning, adaptive routing, congestion control, and fabric monitoring for HDR200/NDR400 GPU clusters.

01

INFINIBAND FABRIC TOPOLOGY DESIGN FOR GPU CLUSTERS

Leaf-spine architecture for up to 256 GPUs: 8 leaf switches (each connecting 8 GPU nodes) interconnected by 2-4 spine switches. Oversubscription ratio must be 1:1 for training clusters to avoid NCCL all-reduce bottlenecks. For 512+ GPUs, three-tier topology with super-spines is required.

Routing algorithm: adaptive routing (AR) distributes traffic across paths based on congestion; deterministic routing (MINIHOP) provides consistent latency. For predictable all-reduce patterns, MINIHOP with 16 LMC per leaf provides best balance of performance and utilization.

Structured cabling: each GPU node has 4-8 HCA ports to different leaf switches for redundancy. Color-coded cables (orange=leaf-spine uplinks, blue=GPU node downlinks, yellow=inter-spine). Labels with source-destination port IDs at both ends.

Cluster SizeTopologyLeaf SwitchesSpine SwitchesRoutingBisection BW
32 GPUsSingle leaf1 x QM97900DirectFull
64 GPUsLeaf-spine 1:11 x QM97901MINIHOP100%
128 GPUsLeaf-spine 1:12 x QM97901-2MINIHOP/AR100%
256 GPUsLeaf-spine 1:14 x QM97902Adaptive100%
512 GPUsThree-tier8 x QM97904 spine+2 superDFSSSP+AR90-100%
1024+ GPUsThree-tier16 x QM97908 spine+4 superUFM adaptive80-100%
02

SUBNET MANAGER CONFIGURATION AND HIGH AVAILABILITY

OpenSM runs on two dedicated management nodes with /etc/opensm/opensm.conf: routing_engine=ftree, sweep_interval=5s. HA: master at sm_priority=0, standby at priority=8. Standby monitors via ibsmvote, promotes self within 15-30s if master fails. Set NCCL_SOCKET_TIMEOUT_MS=60000 to survive failover.

NVIDIA UFM (commercial): web UI, REST API, adaptive routing based on real-time congestion feedback, automated firmware upgrades. Runs on dedicated appliance managing up to 3000+ ports. For clusters >512 GPUs, UFM adaptive routing reduces NCCL latency 10-15% under congestion.

Key OpenSM parameters: routing_engine (ftree for fat-tree, updn for custom), log_file with level 0x02 for errors only, sweep_interval (5s for rapid discovery, 30s for stable fabrics). Active-standby configured via --guid flag pointing to master SM port GUID.

FeatureOpenSMNVIDIA UFMNotes
LicenseOpen Source (BSD)CommercialUFM ~$50-100/port/yr
Routing Algorithmsftree, updn, minhopftree, adaptive, DFSSSPUFM has adaptive routing
HA ModelActive-standby priorityActive-standby + arbitrationBoth 15-30s failover
APICLI + config fileREST API + Python SDKUFM is API-first
Congestion ControlManual static configAuto dynamic feedbackUFM reduces NCCL latency
Max Ports~4000 (limited by sweep)~3000+Both scalable via segmentation
03

PKEY PARTITIONING FOR MULTI-TENANT FABRIC ISOLATION

PKeys provide network-level isolation: 16-bit value with bit 15 indicating full (0x8000|PKey) or limited membership. Assign one partition per tenant: 0x0001 (production), 0x0002 (research), 0x0003 (inference), 0x0004 (storage). Each GPU node can be member of up to 16 partitions.

OpenSM: /etc/opensm/partitions.conf with format partition { name=production pkey=0x7FFF ipoib=ib0.8001 defmember=<GUID> }. UFM: POST /partitions with name, pkey, member GUIDs. ipoib per partition: 0x0001 gets 10.100.1.0/24 on ib0.8001.

PKey enforcement at switch level prevents cross-partition communication. SM computes separate forwarding tables per PKey. NCCL jobs use NCCL_IB_PKEY=0x7FFF to specify partition for communication.

PartitionPKeyMembersIPoIB SubnetTrafficQoS
production-training0x0001Prod GPU nodes10.100.1.0/24NCCL all-reduceHigh (VL15)
research0x0002Research GPU nodes10.100.2.0/24NCCL all-reduceMedium (VL8)
inference0x0003Inference nodes10.100.3.0/24KV cache transferMedium (VL8)
storage-fabric0x0004Storage + GPU nodes10.100.4.0/24Dataset readLow (VL4)
management0x0005BMC + mgmt nodes10.100.5.0/24IPMI/provisioningLowest (VL2)
04

CONGESTION CONTROL AND QUALITY OF SERVICE

Adaptive Routing + Congestion Control: AR selects least-congested paths dynamically, CC uses ECN marking to throttle senders when buffers exceed thresholds. Enable on ConnectX-7: mlxconfig set CC_ENABLE=1, DCQCN_ENABLE=1. Parameters: cc_window=16384, cc_threshold=3, cc_timer=3.0us.

Virtual Lane QoS: VL15 for high-priority NCCL, VL8 for inference, VL2 for storage, VL0 for management. SL2VL mapping: SL0->VL15, SL1->VL8, SL2->VL2. NCCL sets NCCL_IB_SL=0 (production) or NCCL_IB_SL=1 (research).

DCQCN parameters: CLR_TARGET=0.92 (92% link utilization before ECN marking), INIT_ALPHA=8 (slower response, less oscillation), RATE_REDUCE_PERIOD=4us. Proper DCQCN tuning reduces tail latency 20-40% during fabric congestion.

ParameterDefaultGPU Cluster RecommendedEffect
CC_ENABLE0 (disabled)1 (enabled)Enables congestion control
DCQCN_ENABLE0 (disabled)1 (enabled)Enables DCQCN for NCCL
cc_window3256816384Congestion window in 4us units
cc_threshold53Buffer occupancy threshold
cc_timer1.5 us3.0 usSampling interval
CLR_TARGET0.950.92Target utilization before marking
NCCL_IB_SL (prod)00 (VL15)Service level for training
05

FABRIC MONITORING AND TROUBLESHOOTING

Monitoring tools: ibdiagnet --check (daily comprehensive health), iblinkinfo (link state on cabling change), ibqueryerrors (per-port error counters every 5 min to Prometheus), and ibnetdiscover (topology verification). Each feeds into cluster monitoring system.

prometheus-infiniband-exporter: per-port metrics for data rate, xmit_discards, rcv_errors, symbol_errors, link width. Grafana dashboard with fabric utilization heatmap, discard rate for congestion, link state for fault detection.

NCCL debug when performance below baseline: ibdiagnet --check --routing for routing symmetry, ibqueryerrors on switch ports connected to slow nodes, nccl-tests on each node to isolate. If bandwidth <140 GB/s for 8x NDR400, check PKey membership and routing algorithm.

Tool / MetricWhat It RevealsCommandHealthy ThresholdFrequency
ibdiagnet --checkFabric healthibdiagnet --check --extended_statsAll checks passDaily
iblinkinfoLink state/speediblinkinfo | grep -v ActiveAll ACTIVE 4X NDROn cabling change
port_xmit_discardsCongestionibqueryerrors0 per portEvery 5 min
symbol_errorPhysical degradationibqueryerrors< 1/sec per linkEvery 5 min
NCCL all_reduce BWTraining performancenccl-tests all_reduce_perf>160 GB/s (8x NDR400)On job start
Filed under
InfiniBand GPU ClusterNDR400 Fabric SetupInfiniBand PartitioningSubnet Manager GPUOpenSM ConfigurationGPU Cluster NetworkingHDR200 InfiniBand