THE GPU ORCHESTRATOR TRILEMMA
GPU workload orchestration requires choosing between schedulers that were designed for fundamentally different use cases. Kubernetes was designed for stateless microservices and extended for batch GPU scheduling through Volcano and Koordinator. Slurm was designed for HPC batch jobs with fixed resource allocations and gang scheduling. Ray was designed for distributed Python workloads with dynamic task graphs and elastic scaling. Each makes different tradeoffs in scheduling latency (milliseconds vs seconds vs minutes), autoscaling speed (sub-second pod scaling vs per-job queue wait), resource utilization efficiency (bin-packing vs exclusive node allocation vs elastic scaling), and GPU topology awareness.
The practical consequence is that no single orchestrator is optimal for all GPU workloads. Hyperparameter search jobs (thousands of short-lived trials) benefit from Ray's elastic scaling and task-level parallelism. Large-scale LLM training (hundreds of GPUs for weeks) needs Slurm's gang scheduling and deterministic resource allocation. Model inference serving (latency-sensitive, variable traffic) requires Kubernetes with HPA/VPA and GPU-aware request routing. The most sophisticated GPU clusters use all three in a federated orchestration plane, with a common scheduler hierarchy that routes workloads to the appropriate orchestrator based on resource requirements, duration, and SLAs.
| Characteristic | Kubernetes (+Volcano) | Slurm | Ray |
|---|---|---|---|
| Scheduling Unit | Pod (container) | Job (batch step) | Task (Python function) |
| Min Scheduling Latency | 10-50 ms (pod) | 100-500 ms (job) | 1-10 ms (task) |
| Gang Scheduling | Via Volcano / Koordinator | Native (salloc --exclusive) | Via Placement Group |
| GPU Topology Awareness | NVIDIA device plugin + NFD | slurm.conf Gres + topology | Auto-detected by Ray core |
| Autoscaling | Cluster Autoscaler, Karpenter | Manual (no built-in autoscale) | Ray Autoscaler (node-level) |
| Job Prioritization | PriorityClass + Preemption | Priority + Fairshare (multi-factor) | Task queue priorities |
| Multi-Node NCCL Support | Via Kubeflow Training Operator | Native (sbatch -N 8 --gres=gpu:8) | Via Ray Train + NCCL backend |
| Stateful/Inference Serving | Native (Deployment + Service) | Not designed for it | Via Ray Serve |
| Ecosystem Maturity | Broadest (Helm, Istio, Argo) | Deep HPC integrations | Python/ML ecosystem native |
KUBERNETES: MICROSERVICES HERITAGE WITH BATCH GPU EXTENSIONS
Kubernetes handles GPU workloads through the device plugin framework, which advertises `nvidia.com/gpu` resources across 1-8 count depending on GPU density. The vanilla Kubernetes scheduler (kube-scheduler) treats GPUs as discrete countable resources: a pod requesting `nvidia.com/gpu: 8` must land on a node with 8 available GPUs, and the scheduler uses a best-fit algorithm to minimize fragmentation. However, vanilla Kubernetes lacks gang scheduling (all pods in a distributed training job must start simultaneously for NCCL initialization), GPU topology awareness (NVLink domains within a DGX node require GPUs to be scheduled on the same NUMA node), and batch job queue management.
Volcano (formerly kube-batch) addresses these gaps as an alternative Kubernetes scheduler. Volcano's `Queue` and `PodGroup` CRDs implement gang scheduling: all pods in a training job are held until the scheduler can place them simultaneously, then released for parallel startup. Volcano also supports preemption (higher-priority job preempts running pods with PriorityClass), fair-sharing across queues (e.g., research team gets 40 percent of GPU capacity, production gets 60 percent), and resource reservation through the `Reservation` CRD. The `volcano-bj` scheduler plugin adds NUMA-aware GPU placement: when a pod requests 4 GPUs, Volcano ensures they are on the same NVSwitch domain rather than spanning across DGX nodes. The overhead of gang scheduling adds 200-800ms to pod scheduling time but prevents a 30-second NCCL timeout that would occur when 7 of 8 pods start and the 8th is delayed.
| K8s GPU Feature | Vanilla Kubernetes | With Volcano Scheduler | With Koordinator |
|---|---|---|---|
| GPU Scheduling | Fit (first available node) | Gang + Fair-share queue | Elastic bin-packing + NUMA |
| NCCL Multi-Node | Manual (Kubeflow Training Op) | PyTorchJob CRD + gang schedule | Same as Volcano |
| Preemption | PriorityClass (preempt lower priority) | Queue-based preemption + SLAs | Colocation with CPU/memory tiers |
| GPU Oversubscription | Not supported | Not supported | Shared GPU + reclaimable resources |
| Autoscaling GPU nodes | Cluster Autoscaler (+15-60 sec) | Volcano + CA (gang-aware) | Koordinator + CA (colocation-aware) |
| Maturity / Production Readiness | Proven for inference | Proven (Ant Group, Huawei, others) | Proven (Alibaba, 10k+ node clusters) |
SLURM: HPC-GRADE GPU JOB SCHEDULING WITH DETERMINISTIC ALLOCATION
Slurm's design for HPC workloads aligns well with large-scale GPU training: jobs declare resource requirements upfront (`--gres=gpu:8 --nodes=8`), Slurm's backfill scheduler finds placements in the queue that maximize utilization without starving large jobs, and allocations are exclusive and deterministic. The slurm.conf GPU configuration specifies both GRES (Generic Resource Scheduling) for GPU counting and GPU topology via `GresTypes=gpu` and `NodeName=DGX-01 Gres=gpu:8`. Slurm's topology plugin (`TopologyPlugin=topology/tree`) ensures that multi-node GPU jobs are allocated within the same leaf switch to minimize NCCL latency.
Slurm's weakness for GPU clusters is threefold. First, Slurm has no built-in autoscaling: if the cluster has 128 GPUs and a 256-GPU job is submitted, the job stays in the queue indefinitely. GPU clusters using Slurm must integrate with a separate auto-scaling layer (e.g., AWS ParallelCluster or GCP Slurm-GCP) that provisions or deprovisions node groups based on queue depth. Second, Slurm jobs are monolithic: dynamic task spawning (Ray's core use case) requires `srun` within the job allocation, which adds 200-400ms per `srun` call. Third, inference serving is not supported: there is no equivalent to Kubernetes Service for routing requests to GPU pods. Despite these gaps, Slurm remains the dominant orchestrator for clusters where training is the sole or primary workload type, with approximately 40 percent of large >1000-GPU clusters using Slurm in 2026.
RAY: DISTRIBUTED PYTHON WITH ELASTIC GPU EXECUTION
Ray's task-level parallelism is uniquely suited to GPU workloads that require dynamic execution patterns. Unlike Slurm's fixed allocation or Kubernetes' pod-level scheduling, Ray's scheduler (GCS) allocates GPU resources at the task level from a shared pool of Ray nodes. A root task can spawn 64 child tasks that each request 1 GPU, execute for 10 seconds, and terminate, with the GPUs immediately available for the next wave. This is ideal for hyperparameter optimization (Optuna + Ray Tune), reinforcement learning (multiple environment rollout workers each consuming a GPU for a few seconds), and data preprocessing pipelines that use GPU-accelerated libraries like RAPIDS cuDF.
Ray's elastic autoscaler manages GPU node lifecycle based on task demand: `ray up` creates the initial cluster, and the autoscaler adds or removes nodes based on pending GPU task count. The autoscaler configuration specifies `min_workers`, `max_workers`, and `target_utilization_fraction`. For GPU clusters, `target_utilization_fraction: 0.8` triggers new node provisioning when GPU utilization exceeds 80 percent across the cluster. The autoscaler supports spot/preemptible GPU instance integration with `max_failures` and node provider rate limiting to prevent throttling. Ray's weakness is the lack of gang scheduling: if a training job requires 64 GPUs, Ray assigns them on a first-come-first-served basis, potentially allocating 8 GPUs on 8 different nodes when a single 8-GPU DGX node would be more efficient. Ray Placement Groups (`ray.util.placement_group`) provide a workaround by reserving GPUs before task execution, adding 100-500ms reservation latency.
| Workload Type | Best Orchestrator | GPU Count Range | Key Metric |
|---|---|---|---|
| LLM Training (large-scale) | Slurm | 128-4096 | Time-to-train, allocation determinism |
| Hyperparameter Search (1000s of trials) | Ray Tune | 8-256 | Trials/hour, GPU utilization |
| Model Inference Serving | Kubernetes | 1-512 | p99 latency, tokens/sec |
| Reinforcement Learning (parallel envs) | Ray RLlib | 16-128 | Env steps/second, GPU utilization |
| Data Preprocessing (GPU-accelerated) | Ray Data / K8s | 8-64 | Throughput GB/s per GPU |
| Hybrid (training + inference on same cluster) | Kubernetes + Ray | 64-1024 | Resource utilization, scheduling efficiency |
FEDERATED ORCHESTRATION: HYBRID DEPLOYMENT PATTERNS
Real-world GPU clusters above 512 GPUs rarely use a single orchestrator. The federated orchestration pattern uses Kubernetes as the control plane with Slurm and Ray as co-schedulers for batch workloads. The Kubernetes management cluster runs both Slurm controller and Ray head as stateful workloads, with the GPU nodes partitioned into Slurm and Ray node pools via node labels and taints. A unified admission controller (`gpu-scheduler-admission`) intercepts pod creation and routes the pod to the appropriate scheduler based on annotations: `scheduler.clusterbid.com/scheduler: volcano` for batch training, `scheduler.clusterbid.com/scheduler: ray` for distributed Python tasks, and default Kubernetes for inference services.
This federated approach achieves 85-92 percent cluster-wide GPU utilization, versus 60-75 percent for any single scheduler alone. The key enabling technology is the GPU resource pool abstraction: a resource pool of 256 H100 GPUs is dynamically split across orchestrators based on queue depth and priority. Slurm gets 128 GPUs at 9 AM when batch training jobs flood the queue, Ray gets 192 GPUs when hyperparameter sweep jobs run overnight, and Kubernetes inference pods get priority access to 32 GPUs during business hours with burst capacity from the shared pool. The resource pool rebalancing happens through the cluster autoscaler's expander plugins, with a custom `gpu-priority-expander` that selects the node group based on scheduler queue metrics.
