SLURM QOS AND PARTITION DESIGN FOR GPU CLUSTERS
Three QoS levels: high-priority (GrpTRES=gpu=512, MaxWall=168h, Priority=1000), normal (GrpJobs=50/user, GrpTRES=gpu=256, Priority=500), low (GrpTRES=gpu=1024, Priority=100, Preempt=YES). Configured via sacctmgr add qos with respective limits. Partitions map to QoS: gpu-prod, gpu-research, gpu-low.
Physical partitions: gpu-h100-80gb, gpu-a100-80gb, gpu-h100-mig. Logical partitions: gpu-prod (reserved, QoS gpu-prod), gpu-research (shared, QoS gpu-research), gpu-preemptible (shared, QoS gpu-low). Nodes assigned in slurm.conf.
FairShare: each team (account) has FairShare 1-100. Decay half-life 7 days, PriorityWeightAge=0 (no queue time bonus), PriorityFavorSmall=NO. Recalculation every 60s via scontrol reconfig.
Preemption: PreemptMode=REQUEUE, send_user_signal=USR1 at 30s before termination for checkpointing. PreemptType=partition_prio: higher-priority partition preempts lower. USR1 handler in training framework enables seamless resume.
| Parameter | QoS: gpu-prod | QoS: gpu-research | QoS: gpu-low (preempt) |
|---|---|---|---|
| Priority | 1000 | 500 | 100 |
| GrpTRES=gpu | 512 | 256 | 1024 |
| MaxWall | 168:00:00 (7 days) | 24:00:00 | 4:00:00 |
| MaxJobs/user | 10 | 20 | 100 |
| Preempt | No | No | Yes |
| FairShare weight | 100 | 50 | 10 |
| Charge factor | 1.0x | 1.0x | 0.3x |
KUBERNETES RESOURCE QUOTAS AND LIMIT RANGES
ResourceQuota per namespace: spec.hard.nvidia.com/gpu: 32 (GPU limit), requests.nvidia.com/mig-1g.10gb: 16 (MIG slices), limits.cpu: 64, count/deployments.apps: 20. Prevents any namespace from exhausting cluster resources.
LimitRange per namespace: default.nvidia.com/gpu: 2, max.nvidia.com/gpu: 8, min.nvidia.com/gpu: 1. default.cpu: 4, default.memory: 16Gi. Ensures GPU pods have adequate host resources.
PriorityClass: gpu-prod-high (1,000,000), gpu-research-mid (500,000), gpu-development-low (100,000). Higher priority preempts lower when resources are insufficient. PDB protects production inference pods.
| Resource | prod-team ns | research-team ns | ci-cd ns |
|---|---|---|---|
| GPU quota | 64 | 16 | 4 |
| Max GPU per pod | 8 | 4 | 2 |
| Default GPU per pod | 2 | 1 | 1 |
| CPU quota | 128 cores | 32 cores | 8 cores |
| Memory quota | 512 Gi | 128 Gi | 32 Gi |
| Storage quota (SSD) | 10 Ti | 2 Ti | 500 Gi |
| Priority class | prod-high (1M) | research-mid (500K) | dev-low (100K) |
| Preemption allowed | No (PDB protected) | No | Yes |
USER IDENTITY MANAGEMENT AND ACCESS CONTROL
User management integrates with LDAP/AD for centralized identity. SSH key auth via AuthorizedKeysCommand fetching from LDAP. Kubernetes RBAC: ClusterRole gpu-user for submitting pods, gpu-viewer for read-only. OIDC groups mapping.
User onboarding: LDAP account creation, group assignment (gpu-prod-team, gpu-research-team), Slurm account via sacctmgr, Kubernetes RoleBinding. Job submit plugins validate group membership against partition access.
Offboarding: LDAP deactivation, SSH key removal, OIDC token expiry, Slurm MaxJobs=0. Audit log records event. Running jobs terminated via scancel or kubectl delete pods.
PRIORITY SCHEDULING AND PREEMPTION STRATEGIES
Slurm priority: PriorityWeightQOS=1000, PriorityWeightFairShare=500, PriorityWeightPartition=100. QoS is dominant factor, fairshare secondary tiebreaker. PriorityCalcPeriod=60.
Backfill: bf_interval=60, bf_max_time=86400 (schedule up to 24h ahead). Smaller jobs fill gaps left by large jobs waiting for all GPUs. bf_continue keeps backfill running after main scheduling pass.
Gang scheduling (time-slicing): fixed time slices (1h) for interactive workloads. Context switching takes 5-15s for GPU memory load, ~1-3% overhead per hour. Acceptable for development.
Topology-aware placement: select/cons_res + topology/block places jobs on nodes sharing the same leaf switch. topology.conf defines SwitchName hierarchy. Allocates GPUs from minimum switches, improving NCCL bandwidth 15-30%.
| Strategy | Throughput Impact | Fairness | Latency | Best For |
|---|---|---|---|---|
| FIFO | Low (idle GPUs) | High | High | Batch with reservations |
| Backfill | High (fills gaps) | Medium | Low | Mixed 500+ GPUs |
| FairShare (priority) | High | Very high | Medium | Multi-team clusters |
| Preemptive | Very high | Low | Very low | Prod + preemptible mix |
| Gang (time-slice) | Medium (overhead) | High | Very low | Interactive/dev |
| Topology-aware | High (NCCL optimized) | Neutral | Low | Multi-rack clusters |
USAGE MONITORING, REPORTING, AND ENFORCEMENT
Slurm accounting to Elasticsearch via JobCompType=jobcomp/elasticsearch. Daily aggregate per-user GPU-hours, per-partition utilization, queue wait percentiles. Grafana dashboard with per-user self-service view.
Enforcement: job submit plugin checks GPU request against historical utilization. If user averages <40% utilization for requested count over 7 days, plugin suggests reduction. K8s admission webhook validates quota and per-user limits.
Waste tracking: GPU hours allocated but <5% utilized for >30 minutes. Team waste displayed on dashboard, included in weekly report. >20% waste for 2 consecutive weeks: QoS priority reduced 10%. Reduces waste from 15-25% to 8-12% in 3 months.
