WHY TOPOLOGY MATTERS
Standard K8s treats GPUs as fungible. On dual-socket with 8 H100s, GPUs 0-3 connect to NUMA 0, GPUs 4-7 to NUMA 1. Cross-NUMA placement incurs 30-50 percent higher latency and 15-25 percent lower all-reduce performance.
At cluster scale, poor topology placement causes 20-40 percent training throughput degradation. For a $500K training run, this wastes $100-200K.
| Constraint | Impact if Violated | K8s Primitive | Level | Overhead |
|---|---|---|---|---|
| NUMA locality | 15-30% throughput | Topology Manager | Pod | 2-5ms |
| NVLink within node | 5-10% latency | NodeFeatureRule | Node | 10-20ms |
| NVSwitch domain | 20-40% throughput | NodeSelector | Cluster | 50-100ms |
| PCIe generation | 10-15% bandwidth | Node label | Node | Simple |
| Network rack | 10-30% across switches | TopologySpreadConstraints | Cluster | Complex |
NVIDIA GPU OPERATOR AND TOPOLOGY
GPU Operator v23.9+ with NFD labels nodes with nvidia.com/gpu.topology (NVLink connectivity JSON). Pod with annotation nvidia.com/gpu.require-topology=NVSwitch-connected schedules only on fully connected nodes.
kubelet --topology-manager-policy=single-numa-node ensures CPU, memory, and GPU from same NUMA. Use 'restricted' policy for pods needing 8 GPUs across both NUMA nodes.
CUSTOM SCHEDULER EXTENSIONS
Volcano scheduler with gpupack plugin implements bin-packing, increasing GPU utilization by 10-18 percent by preserving contiguous GPU blocks for training jobs.
A 128-node H100 cluster with Volcano achieved: 94 percent GPU utilization (vs 78% default), 22 percent higher throughput, 35 percent fewer OOM failures.
| Scheduler | GPU Utilization | Throughput vs Default | Job Failure Rate | Complexity |
|---|---|---|---|---|
| Default K8s | 78% | Baseline | 8.5% | None |
| GPU Operator + Topo Mgr | 85% | 12% | 5.2% | Low |
| Volcano + gpupack | 94% | 22% | 3.1% | Medium |
| Custom scheduler | 93% | 20% | 2.8% | High |
| Slurm topology-aware | 96% | 25% | 2.1% | High |
NETWORK TOPOLOGY BEYOND SINGLE NODE
Jobs sharing a leaf switch achieve 15-25 percent higher all-reduce bandwidth. Label nodes with rack/switch location using NFD/InfiniBand GUID.
Most clusters use 3:1 or 4:1 oversubscribed leaf-spine. Topology-aware scheduling that co-locates job nodes within a leaf switch reduces training time by 10-20 percent for LLM training.
