All essays
BenchmarkCOMPARISONFEB 2026

Kubernetes GPU Operator: NVIDIA vs AMD Setup, Configuration, and Production Deployment

NVIDIA GPU Operator vs AMD GPU Operator for Kubernetes: architecture comparison, Helm configuration, MIG/SR-IOV device plugin setup, monitoring stack, node feature discovery, and production deployment patterns for H100 and MI400 GPU clusters.

01

GPU OPERATOR ARCHITECTURE: NVIDIA VS AMD

The NVIDIA GPU Operator (v24.9) deploys and manages all NVIDIA GPU software components on Kubernetes through a set of operators: the NVIDIADriver daemon handles driver installation; the NVIDIAContainerToolkit daemon set configures the container runtime; the NVIDIADevicePlugin exposes GPUs as extend resources; the NVIDIAMIGManager partitions GPUs into MIG slices; the GFD (GPU Feature Discovery) labels nodes with GPU attributes; and the DCGMExporter exports GPU metrics to Prometheus. The operator is installed via a single Helm chart: helm install gpu-operator nvidia/gpu-operator --namespace gpu-operator --create-namespace.

The AMD GPU Operator (v2.5) follows a similar pattern but with an important architectural difference: AMD splits its operator into the amd-gpu-operator (which manages the ROCm driver, container runtime, and device plugin) and the k8s-device-plugin (which manages SR-IOV virtual functions). AMD's operator uses Node Feature Discovery (NFD) from the Kubernetes ecosystem rather than a custom feature discovery component. The AMD operator's Helm chart is: helm install amd-gpu-operator amd/amd-gpu-operator --namespace amd-gpu --create-namespace.

ComponentNVIDIA GPU Operator 24.9AMD GPU Operator 2.5
Driver ManagementNVIDIADriver daemon (auto-download)ROCmDriver daemon (kernel module)
Container Runtimenvidia-container-toolkitrocm-container-runtime
Device Pluginnvidia.com/gpu (MIG-aware)amd.com/gpu (SR-IOV-aware)
GPU PartitioningMIGManager (MIG profiles)SR-IOV VF configuration
Feature DiscoveryGPU Feature Discovery (custom)Node Feature Discovery (K8s ecosystem)
Metrics ExportDCGM Exporter (Prometheus)AMDSMI Exporter (Prometheus)
ValidationGPU health validator daemonROCm health check init container
Helm Chart Size200+ configurable values80+ configurable values
02

NVIDIA GPU OPERATOR: PRODUCTION SETUP AND VALUES TUNING

A production NVIDIA GPU Operator deployment requires careful values.yaml configuration. The critical parameters: devicePlugin.enabled=true (default, required for GPU scheduling); migManager.enabled=true for MIG-partitioned GPUs; gfd.enabled=true for GPU feature labels; dcgmExporter.enabled=true for Prometheus metrics; toolkit.enabled=true for container runtime configuration. For H100 clusters, set toolkit.version: v1.18.0-ubuntu24.04 for CUDA 13 compatibility. The migManager.migProfileDefaults must match the GPU SKU: H100 SXM supports profiles all-1g.11gb, all-2g.22gb, all-3g.46gb, all-4g.46gb, all-7g.78gb.

The validation command sequence for a healthy NVIDIA GPU Operator deployment: kubectl get pods -n gpu-operator (all pods Running); kubectl describe node | grep nvidia.com/gpu (GPU capacity listed); kubectl logs -n gpu-operator -l app=nvidia-device-plugin-daemon (device plugin registered each GPU). A common production issue is the driver compatibility check failing on RHEL-based nodes with Secure Boot enabled. Workaround: helm upgrade gpu-operator nvidia/gpu-operator --set driver.mksImage=nvidia/mkisofs:latest --set driver.rdma.enabled=true.

ConfigurationHelm ValueDefaultProduction Recommendation
MIG ManagermigManager.enabledfalsetrue (for MIG deployment)
Default MIG ProfilemigManager.defaultGPUsAsMIGEnabledfalsetrue (if all GPUs use MIG)
Container Runtimetoolkit.versionv1.14.0-ubuntu22.04v1.18.0-ubuntu24.04 (CUDA 13)
MetricsdcgmExporter.enabledtruetrue
Feature Discoverygfd.enabledtruetrue
GPU Resource NamedevicePlugin.resourceNamenvidia.com/gpunvidia.com/gpu (default)
Node Feature Labelsgfd.plugins[gpu, mig, ...][gpu, mig, memory, timeout]
03

AMD GPU OPERATOR: SR-IOV CONFIGURATION AND ROCM DEPLOYMENT

The AMD GPU Operator configuration centers on SR-IOV virtual function management. The key values.yaml parameters: devicePlugin.enabled=true; sriovDevicePlugin.enabled=true (SR-IOV VF allocation); rocmDriver.enabled=true (ROCm kernel module installation); rocmDriver.version: 6.5.0; amdsmiExporter.enabled=true. The SR-IOV configuration specifies VF count per GPU: sriovDevicePlugin.sriovResources[0].vfAmount: 8 creates 8 virtual functions per MI400 GPU. AMD's device plugin registers each VF as a separate amd.com/gpu resource.

Production AMD GPU Operator validation: kubectl get pods -n amd-gpu (all pods Running); kubectl exec -n amd-gpu deploy/amd-smi-exporter -- rocm-smi --showvf (lists VF topology). AMD's operator supports dynamic SR-IOV reconfiguration: kubectl annotate node <node> amd.com/sriov-vf-count=12 changes VF count without draining workloads. The RuntimeClass for AMD GPU pods must be configured as runtimeClassName: amd-gpu in the pod spec.

04

MONITORING, ALERTING, AND OBSERVABILITY

Both operators export GPU metrics to Prometheus via their respective exporters. NVIDIA's DCGM Exporter provides 250+ GPU metrics including: DCGM_FI_DEV_GPU_UTIL (GPU utilization 0-100%), DCGM_FI_DEV_MEM_COPY_UTIL (memory bandwidth utilization), DCGM_FI_DEV_POWER_USAGE (power draw in milliwatts), DCGM_FI_DEV_SM_CLOCK (SM clock frequency), DCGM_FI_DEV_NVLINK_BANDWIDTH_TOTAL (total NVLink bandwidth), and DCGM_FI_DEV_GPU_TEMP (temperature in Celsius). The recommended Prometheus recording rules include avg by (node) (rate(DCGM_FI_DEV_GPU_UTIL[5m])) for cluster GPU utilization.

AMD's AMDSMI Exporter provides comparable metrics: amdgpu_utilization_percent, amdgpu_memory_used_bytes, amdgpu_power_draw_watts, amdgpu_temperature_celsius, and amdgpu_vf_count (number of active VFs). AMD's metrics coverage is narrower than NVIDIA's: 80 metrics versus 250, with no SM occupancy or memory bandwidth utilization metrics. The recommended alert thresholds for both operators: GPU utilization < 30% for > 5 minutes, GPU temperature > 85C, GPU power > 95% of TDP for > 10 minutes, and memory utilization > 90%.

Monitoring DimensionNVIDIA DCGM ExporterAMD AMDSMI Exporter
Total Metrics Exported250+80+
GPU UtilizationYes (per-GPU)Yes (per-VF)
Memory UtilizationYes (allocated/free/total)Yes (used/total)
Power DrawYes (mW)Yes (W)
TemperatureYes (multiple sensors)Yes (junction/memory)
NVLink BandwidthYes (per-link)N/A (Infinity Fabric)
SM OccupancyYesNot available
Memory Bandwidth UtilizationYesNot available
VF Count / MIG SlicesYes (per-profile)Yes (per-VF)
05

PRODUCTION DEPLOYMENT PATTERNS AND COMMON PITFALLS

For multi-node GPU training on Kubernetes, both operators support node-affinity-based GPU scheduling. The pattern: label GPU nodes with gpu-type: h100 or gpu-type: mi400, then use nodeSelector and tolerations in pod specs. For NVLink-connected GPU pods, set requiredDuringSchedulingIgnoredDuringExecution node affinity to ensure all GPUs are on the same node. For multi-node training with InfiniBand, both operators support RDMA device plugin registration.

Common pitfalls: First, MIG configuration ordering - NVIDIA GPU Operator requires MIG to be configured BEFORE the device plugin starts, otherwise GPU capacity is double-counted. Fix: helm upgrade --set migManager.enabled=true --set devicePlugin.enabled=false && helm upgrade --set devicePlugin.enabled=true in two steps. Second, AMD SR-IOV VF exhaustion - each VF consumes 1 GB of host memory for DMA buffers; 16 VFs per MI400 on a 32-GPU node consumes 512 GB host RAM. Third, GPU time-slicing conflicts with MIG: if both timeSlicing.enabled and migManager.enabled are true, the time-slicing scheduler can preempt MIG slices.

06

GPU OPERATOR READY INSTANCES ON CLUSTERBID

ClusterBid's provider network includes GPU instances pre-configured with both NVIDIA GPU Operator and AMD GPU Operator support. The --gpu-operator filter identifies instances with the operator pre-installed or available for installation. NVIDIA GPU Operator-ready instances across 12 providers start at $3.10/hour for a 1x H100 node and $28.00/hour for an 8x H100 NVLink node. AMD GPU Operator-ready MI400 instances start at $2.40/hour for a 1x MI400 node.

The recommended ClusterBid workflow for GPU Operator deployment: (1) Filter by GPU operator readiness and GPU type. (2) Select a provider with documented Helm values. (3) Use ClusterBid's --setup-script feature to automate helm install as a post-provisioning step. (4) Validate with the gpu-operator-validator job. (5) Deploy AI workloads with kubectl apply -f training-job.yaml using nvidia.com/gpu: 8 or amd.com/gpu: 8 resource requests.

Filed under
Kubernetes GPU OperatorNVIDIA GPU OperatorAMD GPU OperatorK8s GPU Device PluginMIG KubernetesSR-IOV KubernetesGPU Cluster Helm