All essays
InfrastructureINFRASTRUCTUREFEB 2026

GPU Cluster Topology-Aware Scheduling for Kubernetes

Topology-aware GPU scheduling for Kubernetes. Schedule pods with NUMA affinity, NVLink topology, and GPU bandwidth awareness for optimal H100 cluster performance.

01

WHY TOPOLOGY MATTERS

Standard K8s treats GPUs as fungible. On dual-socket with 8 H100s, GPUs 0-3 connect to NUMA 0, GPUs 4-7 to NUMA 1. Cross-NUMA placement incurs 30-50 percent higher latency and 15-25 percent lower all-reduce performance.

At cluster scale, poor topology placement causes 20-40 percent training throughput degradation. For a $500K training run, this wastes $100-200K.

ConstraintImpact if ViolatedK8s PrimitiveLevelOverhead
NUMA locality15-30% throughputTopology ManagerPod2-5ms
NVLink within node5-10% latencyNodeFeatureRuleNode10-20ms
NVSwitch domain20-40% throughputNodeSelectorCluster50-100ms
PCIe generation10-15% bandwidthNode labelNodeSimple
Network rack10-30% across switchesTopologySpreadConstraintsClusterComplex
02

NVIDIA GPU OPERATOR AND TOPOLOGY

GPU Operator v23.9+ with NFD labels nodes with nvidia.com/gpu.topology (NVLink connectivity JSON). Pod with annotation nvidia.com/gpu.require-topology=NVSwitch-connected schedules only on fully connected nodes.

kubelet --topology-manager-policy=single-numa-node ensures CPU, memory, and GPU from same NUMA. Use 'restricted' policy for pods needing 8 GPUs across both NUMA nodes.

03

CUSTOM SCHEDULER EXTENSIONS

Volcano scheduler with gpupack plugin implements bin-packing, increasing GPU utilization by 10-18 percent by preserving contiguous GPU blocks for training jobs.

A 128-node H100 cluster with Volcano achieved: 94 percent GPU utilization (vs 78% default), 22 percent higher throughput, 35 percent fewer OOM failures.

SchedulerGPU UtilizationThroughput vs DefaultJob Failure RateComplexity
Default K8s78%Baseline8.5%None
GPU Operator + Topo Mgr85%12%5.2%Low
Volcano + gpupack94%22%3.1%Medium
Custom scheduler93%20%2.8%High
Slurm topology-aware96%25%2.1%High
04

NETWORK TOPOLOGY BEYOND SINGLE NODE

Jobs sharing a leaf switch achieve 15-25 percent higher all-reduce bandwidth. Label nodes with rack/switch location using NFD/InfiniBand GUID.

Most clusters use 3:1 or 4:1 oversubscribed leaf-spine. Topology-aware scheduling that co-locates job nodes within a leaf switch reduces training time by 10-20 percent for LLM training.

Filed under
Kubernetes GPUTopology-Aware SchedulingNUMA AffinityNVLink TopologyGPU OperatorNode Feature Discovery