INFINIBAND FABRIC TOPOLOGY DESIGN FOR GPU CLUSTERS
Leaf-spine architecture for up to 256 GPUs: 8 leaf switches (each connecting 8 GPU nodes) interconnected by 2-4 spine switches. Oversubscription ratio must be 1:1 for training clusters to avoid NCCL all-reduce bottlenecks. For 512+ GPUs, three-tier topology with super-spines is required.
Routing algorithm: adaptive routing (AR) distributes traffic across paths based on congestion; deterministic routing (MINIHOP) provides consistent latency. For predictable all-reduce patterns, MINIHOP with 16 LMC per leaf provides best balance of performance and utilization.
Structured cabling: each GPU node has 4-8 HCA ports to different leaf switches for redundancy. Color-coded cables (orange=leaf-spine uplinks, blue=GPU node downlinks, yellow=inter-spine). Labels with source-destination port IDs at both ends.
| Cluster Size | Topology | Leaf Switches | Spine Switches | Routing | Bisection BW |
|---|---|---|---|---|---|
| 32 GPUs | Single leaf | 1 x QM9790 | 0 | Direct | Full |
| 64 GPUs | Leaf-spine 1:1 | 1 x QM9790 | 1 | MINIHOP | 100% |
| 128 GPUs | Leaf-spine 1:1 | 2 x QM9790 | 1-2 | MINIHOP/AR | 100% |
| 256 GPUs | Leaf-spine 1:1 | 4 x QM9790 | 2 | Adaptive | 100% |
| 512 GPUs | Three-tier | 8 x QM9790 | 4 spine+2 super | DFSSSP+AR | 90-100% |
| 1024+ GPUs | Three-tier | 16 x QM9790 | 8 spine+4 super | UFM adaptive | 80-100% |
SUBNET MANAGER CONFIGURATION AND HIGH AVAILABILITY
OpenSM runs on two dedicated management nodes with /etc/opensm/opensm.conf: routing_engine=ftree, sweep_interval=5s. HA: master at sm_priority=0, standby at priority=8. Standby monitors via ibsmvote, promotes self within 15-30s if master fails. Set NCCL_SOCKET_TIMEOUT_MS=60000 to survive failover.
NVIDIA UFM (commercial): web UI, REST API, adaptive routing based on real-time congestion feedback, automated firmware upgrades. Runs on dedicated appliance managing up to 3000+ ports. For clusters >512 GPUs, UFM adaptive routing reduces NCCL latency 10-15% under congestion.
Key OpenSM parameters: routing_engine (ftree for fat-tree, updn for custom), log_file with level 0x02 for errors only, sweep_interval (5s for rapid discovery, 30s for stable fabrics). Active-standby configured via --guid flag pointing to master SM port GUID.
| Feature | OpenSM | NVIDIA UFM | Notes |
|---|---|---|---|
| License | Open Source (BSD) | Commercial | UFM ~$50-100/port/yr |
| Routing Algorithms | ftree, updn, minhop | ftree, adaptive, DFSSSP | UFM has adaptive routing |
| HA Model | Active-standby priority | Active-standby + arbitration | Both 15-30s failover |
| API | CLI + config file | REST API + Python SDK | UFM is API-first |
| Congestion Control | Manual static config | Auto dynamic feedback | UFM reduces NCCL latency |
| Max Ports | ~4000 (limited by sweep) | ~3000+ | Both scalable via segmentation |
PKEY PARTITIONING FOR MULTI-TENANT FABRIC ISOLATION
PKeys provide network-level isolation: 16-bit value with bit 15 indicating full (0x8000|PKey) or limited membership. Assign one partition per tenant: 0x0001 (production), 0x0002 (research), 0x0003 (inference), 0x0004 (storage). Each GPU node can be member of up to 16 partitions.
OpenSM: /etc/opensm/partitions.conf with format partition { name=production pkey=0x7FFF ipoib=ib0.8001 defmember=<GUID> }. UFM: POST /partitions with name, pkey, member GUIDs. ipoib per partition: 0x0001 gets 10.100.1.0/24 on ib0.8001.
PKey enforcement at switch level prevents cross-partition communication. SM computes separate forwarding tables per PKey. NCCL jobs use NCCL_IB_PKEY=0x7FFF to specify partition for communication.
| Partition | PKey | Members | IPoIB Subnet | Traffic | QoS |
|---|---|---|---|---|---|
| production-training | 0x0001 | Prod GPU nodes | 10.100.1.0/24 | NCCL all-reduce | High (VL15) |
| research | 0x0002 | Research GPU nodes | 10.100.2.0/24 | NCCL all-reduce | Medium (VL8) |
| inference | 0x0003 | Inference nodes | 10.100.3.0/24 | KV cache transfer | Medium (VL8) |
| storage-fabric | 0x0004 | Storage + GPU nodes | 10.100.4.0/24 | Dataset read | Low (VL4) |
| management | 0x0005 | BMC + mgmt nodes | 10.100.5.0/24 | IPMI/provisioning | Lowest (VL2) |
CONGESTION CONTROL AND QUALITY OF SERVICE
Adaptive Routing + Congestion Control: AR selects least-congested paths dynamically, CC uses ECN marking to throttle senders when buffers exceed thresholds. Enable on ConnectX-7: mlxconfig set CC_ENABLE=1, DCQCN_ENABLE=1. Parameters: cc_window=16384, cc_threshold=3, cc_timer=3.0us.
Virtual Lane QoS: VL15 for high-priority NCCL, VL8 for inference, VL2 for storage, VL0 for management. SL2VL mapping: SL0->VL15, SL1->VL8, SL2->VL2. NCCL sets NCCL_IB_SL=0 (production) or NCCL_IB_SL=1 (research).
DCQCN parameters: CLR_TARGET=0.92 (92% link utilization before ECN marking), INIT_ALPHA=8 (slower response, less oscillation), RATE_REDUCE_PERIOD=4us. Proper DCQCN tuning reduces tail latency 20-40% during fabric congestion.
| Parameter | Default | GPU Cluster Recommended | Effect |
|---|---|---|---|
| CC_ENABLE | 0 (disabled) | 1 (enabled) | Enables congestion control |
| DCQCN_ENABLE | 0 (disabled) | 1 (enabled) | Enables DCQCN for NCCL |
| cc_window | 32568 | 16384 | Congestion window in 4us units |
| cc_threshold | 5 | 3 | Buffer occupancy threshold |
| cc_timer | 1.5 us | 3.0 us | Sampling interval |
| CLR_TARGET | 0.95 | 0.92 | Target utilization before marking |
| NCCL_IB_SL (prod) | 0 | 0 (VL15) | Service level for training |
FABRIC MONITORING AND TROUBLESHOOTING
Monitoring tools: ibdiagnet --check (daily comprehensive health), iblinkinfo (link state on cabling change), ibqueryerrors (per-port error counters every 5 min to Prometheus), and ibnetdiscover (topology verification). Each feeds into cluster monitoring system.
prometheus-infiniband-exporter: per-port metrics for data rate, xmit_discards, rcv_errors, symbol_errors, link width. Grafana dashboard with fabric utilization heatmap, discard rate for congestion, link state for fault detection.
NCCL debug when performance below baseline: ibdiagnet --check --routing for routing symmetry, ibqueryerrors on switch ports connected to slow nodes, nccl-tests on each node to isolate. If bandwidth <140 GB/s for 8x NDR400, check PKey membership and routing algorithm.
| Tool / Metric | What It Reveals | Command | Healthy Threshold | Frequency |
|---|---|---|---|---|
| ibdiagnet --check | Fabric health | ibdiagnet --check --extended_stats | All checks pass | Daily |
| iblinkinfo | Link state/speed | iblinkinfo | grep -v Active | All ACTIVE 4X NDR | On cabling change |
| port_xmit_discards | Congestion | ibqueryerrors | 0 per port | Every 5 min |
| symbol_error | Physical degradation | ibqueryerrors | < 1/sec per link | Every 5 min |
| NCCL all_reduce BW | Training performance | nccl-tests all_reduce_perf | >160 GB/s (8x NDR400) | On job start |
