GRES GPU CONFIGURATION IN SLURM
Slurm manages GPU resources through the GRES (Generic Resource Scheduling) plugin. The slurm.conf configuration defines GPU resources at the node level: GresTypes=gpu and per-node NodeName=h100-gpu-[1-32] Gres=gpu:8 for 8-GPU nodes. The gres.conf file specifies GPU device files and MIG configurations: Name=gpu Type=h100 File=/dev/nvidia[0-7]. For MIG-partitioned GPUs, each MIG slice is defined as a separate GRES: Name=gpu Type=h100-mig-1g File=/dev/nvidia-caps/nvidia-cap[0-55]. NVIDIA provides the slurm-gres-nvidia-plugin which auto-detects GPU topology and MIG profiles.
The GRES plugin supports GPU affinity through --gres-flags=enforce-binding which pins jobs to specific GPU devices. On H100 nodes, --gres=gpu:4 combined with --gres-flags=enforce-binding allocates 4 specific GPUs and sets CUDA_VISIBLE_DEVICES to those device IDs. Slurm 24.11+ adds GPU topology-aware allocation: --gres-flags=topology_fit places GPU allocations within the same NVSwitch domain when possible, significantly reducing all-reduce latency.
GPU PARTITION DESIGN AND NODE GROUPING
Production Slurm GPU clusters organize nodes into partitions based on GPU type, interconnect, and workload characteristics. A typical H100 cluster partition scheme: training-h100 (8xH100 NVLink nodes, full-node jobs only, InfiniBand interconnect), inference-h100 (8xH100 PCIe nodes, fractional GPU jobs, Ethernet interconnect), mig-h100 (H100 nodes with MIG 1g slices, single-slice containers), and interactive-h100 (2xH100 nodes, short time limit, low priority). For mixed-vendor clusters, separate partitions prevent workload placement on wrong GPU architectures.
Partition configuration in slurm.conf: PartitionName=training-h100 Nodes=h100-trn-[1-128] Default=NO MaxNodes=128 DefaultTime=24:00:00 State=UP. The OverSubscribe=NO directive prevents CPU overcommit on GPU partitions. AMD MI400 nodes use a separate partition with the rocm feature label: PartitionName=amd-mi400 Nodes=mi400-[1-32] Feature=rocm. The Priority parameter controls partition-level scheduling priority.
| Partition | Node Type | GPU Count | Interconnect | Job Type | Default Time Limit |
|---|---|---|---|---|---|
| training-h100 | H100 SXM | 8 per node | NVSwitch + InfiniBand NDR400 | Full-node training jobs | 24 hours |
| inference-h100 | H100 PCIe | 8 per node | PCIe Gen5 + Ethernet 100G | Fractional GPU serving | 7 days |
| mig-h100 | H100 MIG-enabled | 56 MIG slices | NVLink (shared per GPU) | Single-slice inference | 48 hours |
| b200-training | B200 SXM | 8 per node | NVSwitch 5 + InfiniBand XDR | Full-node training | 24 hours |
| amd-mi400 | MI400 | 8 per node | Infinity Fabric + IB NDR400 | ROCm training | 24 hours |
| interactive | H100 SXM | 2 per node | Ethernet 25G | Dev/debug sessions | 1 hour |
| debug | H100 SXM | 8 per node | NVSwitch | Quick testing | 1 hour |
JOB SCHEDULING POLICIES FOR GPU WORKLOADS
Slurm's scheduling policies significantly impact GPU cluster utilization. The default FIFO scheduler leaves GPUs idle when large jobs queue behind smaller ones. Backfill scheduling (SchedulerType=sched/backfill) enables small jobs to run while large jobs wait for resources, improving utilization by 15-30% on GPU clusters. For training jobs requiring 8 or more GPUs, gang scheduling (GangSchedType=preempt) allows higher-priority jobs to preempt running jobs with checkpoint resumption.
GPU topology-aware scheduling in Slurm 24.11+ uses the TopologyParam=TopologyTree directive to build a tree of GPU interconnects. Jobs requesting --switches=1@00:10:00 require all allocated GPUs to be within one switch domain. The --gpu-bind=none|per_task|closest flag controls GPU-to-CPU binding. For AI training, --gpu-bind=closest binds each GPU to the nearest NUMA domain, reducing PCIe latency by 8-12%.
| Scheduler Configuration | Parameter | GPU Cluster Impact | Recommendation |
|---|---|---|---|
| Backfill Scheduling | sched/backfill | +15-30% utilization | Enable bf_interval=60 |
| Gang Scheduling | preempt-type=gang | +20% throughput for mixed jobs | Enable with checkpoint |
| Topology-Aware | TopologyTree + switches | -25% all-reduce latency | Required for NVSwitch |
| Fairshare | PriorityType=priority/multifactor | Prevents team resource hogging | Enable with 3 factors |
| GPU Exclusive | --exclusive | Prevents GPU sharing conflicts | Default for training |
| Job Array | --array=1-100%4 | Parallel hyperparameter tuning | Limit concurrent %N |
GPU ACCOUNTING AND FAIRSHARE
Slurm's accounting infrastructure tracks GPU usage through the sacct and sacctmgr commands. GPU accounting requires AccountingStorageTRES=gpu in slurm.conf and the jobacct_gather/cgroup plugin to report per-job GPU utilization. The sacct --format=JobID,Elapsed,AllocTRES%40,ConsumedEnergyRaw,GPUFrac command shows per-job GPU allocation and utilization fraction. The GPUFrac field, new in Slurm 24.11, reports actual GPU utilization as a fraction of allocated GPU time.
Fairshare scheduling (PriorityType=priority/multifactor) for GPU clusters uses three factors: Fairshare (usage history), Partition Priority (partition-level weighting), and Job Size (large GPU jobs get higher priority). For GPU clusters shared between training and inference teams, configure separate associations with different fairshare weights. GPU chargeback tracks GPU-hours consumed per team via sacctmgr, supporting internal billing at $2.50-5.00/GPU-hour for H100 resources.
PRODUCTION SLURM.CONF FOR GPU CLUSTERS
A production slurm.conf for a 128-node H100 GPU cluster includes these critical directives: GresTypes=gpu,mps (enable GPU and MPS sharing), SelectType=select/cons_res (consumable resource selection), SelectTypeParameters=CR_GPU (GPU as a consumable resource), PriorityType=priority/multifactor (fairshare), SchedulerType=sched/backfill (backfill scheduling), PreemptMode=gang (gang preemption), TaskPlugin=task/affinity,task/cgroup (GPU binding), PrologFlags=Contain (containerize prolog/epilog).
Performance-critical Slurm configuration parameters for GPU workloads: MessageTimeout=60 (reduce from default 120 for faster job startup on InfiniBand clusters), TreeWidth=512 (optimized for 128+ node GPU clusters), SchedulerParameters=default_queue_depth=10000,sched_interval=30,bf_interval=60,bf_max_time=60000,bf_resolution=500. The bf_resolution=500 millisecond setting enables sub-second GPU allocation granularity for backfill scheduling. SchedulerParameters=max_gpu_job_launch=8 limits concurrent GPU job launches.
SLURM-READY GPU INSTANCES ON CLUSTERBID
ClusterBid lists 3,200+ GPU instances across 14 providers that support Slurm workload manager. The --scheduler slurm filter identifies providers offering pre-installed Slurm or Slurm-on-demand installation. Instances are tagged with Slurm version (slurm-24.11, slurm-24.08), GRES configuration (gres/gpu:8), and MIG support. For dedicated Slurm GPU clusters, ClusterBid partners with 6 providers offering managed Slurm control nodes with automatic partition configuration.
Pricing for Slurm-managed GPU clusters on ClusterBid: 8xH100 Slurm partition at $24-36/hour; 8xMI400 Slurm partition at $18-28/hour; 8xB200 Slurm partition at $40-60/hour. The --slurm-config parameter accepts a custom slurm.conf fragment that ClusterBid applies during provisioning. Slurm-enabled H100 partitions deliver 92-95% scheduling utilization versus 70-80% for shared cloud GPU queues, with 30-45% lower effective cost per GPU-hour.
