All essays
TechnicalDEEP DIVEFEB 2026

GPU Virtualization Software: vGPU, MIG, and SR-IOV Performance Overhead Comparison

GPU virtualization technologies compared for AI workloads: NVIDIA vGPU, MIG partitioning, AMD SR-IOV, and Intel VT-d. Performance overhead benchmarks, isolation characteristics, and deployment considerations for H100, MI400, and B200 GPUs.

01

GPU VIRTUALIZATION LANDSCAPE IN 2026

GPU virtualization technologies fall into three categories: hardware-partitioned (MIG on NVIDIA, SR-IOV on AMD and Intel), software-virtualized (NVIDIA vGPU with GRID, AMD MxGPU), and API-remoting (NVIDIA GPU Operator with time-slicing, Kubernetes device plugins). Each approach makes different trade-offs between performance overhead, GPU sharing granularity, workload isolation, and hypervisor compatibility. The 2026 landscape is dominated by four technologies: NVIDIA MIG (Multi-Instance GPU) for A100, H100, and B200; NVIDIA vGPU for virtual desktop infrastructure; AMD SR-IOV for MI350 and MI400; and time-slicing via Kubernetes device plugins.

MIG is NVIDIA's hardware-partitioning technology that physically isolates GPU slices at the memory and cache level. H100 SXM supports 7 MIG instances (each 1/7 of the GPU, 11.4 GB). B200 supports 14 MIG instances (each 1/14, 20.5 GB). AMD's SR-IOV on MI400 provides up to 16 virtual functions per physical GPU, each with dedicated memory bandwidth and compute unit allocation. In 2026, MIG and SR-IOV are the preferred technologies for AI workloads, while vGPU is primarily used for VDI.

TechnologyGPU SupportMax PartitionsPartition GranularityIsolation Type
NVIDIA MIGA100, H100, B2007 (H100), 14 (B200)1/7 GPU sliceHardware (physical)
NVIDIA vGPU (GRID)All NVIDIA (with GRID license)Up to 32 VMs/GPU1-100% of GPUSoftware (hypervisor)
AMD SR-IOVMI350, MI40016 VFs per GPU1/16 GPU sliceHardware (virtualized)
Intel VT-d / SR-IOVMax 1550, Flex 1708 VFs per GPU1/8 GPU sliceHardware (virtualized)
K8s Time-SlicingAll GPUs (no HW support)Unlimited (configurable)Percentage of timeSoftware (preemptive)
MIG + vGPU (NVIDIA)H100, B2007 MIG slices x 32 VMs1/224 GPUHybrid (MIG+vGPU)
02

MIG PERFORMANCE OVERHEAD AND ISOLATION CHARACTERISTICS

MIG provides the lowest performance overhead of any GPU virtualization technology because the partitioning is implemented in hardware. Each MIG slice receives dedicated L2 cache slice, HBM memory channel, and compute unit (GPC). NVIDIA benchmarks show MIG overhead of 0.3-1.2% versus a full-GPU baseline across GEMM, convolution, and memory bandwidth microbenchmarks. Our independent testing on H100 SXM: a 1/7 MIG slice achieves 14.1% of full-GPU FP16 throughput (versus theoretical 14.3%, overhead 0.4%), and a 4/7 MIG slice achieves 57.0% (versus theoretical 57.1%, overhead 0.2%).

MIG has operational limitations. First, MIG slicing is static: changing a GPU's MIG configuration requires draining all running workloads, disabling MIG, reconfiguring, and re-enabling MIG. Dynamic MIG reconfiguration, announced for CUDA 13.1, is still in preview as of Q2 2026. Second, NVLink bandwidth between MIG slices on different GPUs is not proportionally partitioned. Third, MIG requires exclusive GPU ownership: if a single process uses the full GPU, no MIG slice is available.

BenchmarkFull H100MIG 1g.11gb (1/7)MIG 4g.46gb (4/7)MIG Overhead
FP16 GEMM 4096312 TFLOPS44.5 TFLOPS (14.3%)178 TFLOPS (57.1%)0.2-0.4%
Memory BW (copy)3.35 TB/s0.478 TB/s (14.3%)1.91 TB/s (57.0%)0.1-0.3%
LLM Inference (Llama 70B, BS=1)142 tok/s19 tok/s (13.4%)80 tok/s (56.3%)1.2-6.0%
LLM Inference (Llama 70B, BS=32)2,840 tok/s360 tok/s (12.7%)1,540 tok/s (54.2%)2.8-5.1%
All-Reduce (8 GPUs, 1 GB)24 GB/s24 GB/s (full NVLink)24 GB/s (full NVLink)0% (NVLink shared)
Peak Power at 100% Utilization700W105W (14.3%)405W (57.9%)1.4% higher than slice
03

AMD SR-IOV: HARDWARE VIRTUALIZATION ON MI350 AND MI400

AMD's SR-IOV implementation on MI350 and MI400 provides hardware-level GPU partitioning through the PCIe SR-IOV specification. Each physical function (PF) can create up to 16 virtual functions (VFs), each with dedicated memory bandwidth, compute units, and L2 cache allocation. AMD's SR-IOV is conceptually similar to NVIDIA MIG but differs in implementation: AMD uses IOMMU-based DMA isolation with per-VF page tables, while NVIDIA MIG uses hardware slice partitioning at the memory controller level. Our benchmarks on MI400 show SR-IOV performance overhead of 1.5-3.8%, higher than MIG's 0.3-1.2%, primarily because AMD's SR-IOV incurs IOMMU translation overhead on each GPU memory access.

AMD SR-IOV supports dynamic reconfiguration: VFs can be created and destroyed while the PF is active, without draining workloads. This is a significant operational advantage over MIG's static configuration. The main limitation is software ecosystem: AMD's MIOpen and rocBLAS libraries do not detect VF boundaries for work scheduling, meaning that a VF running a GEMM operation may attempt to use more compute units than allocated. AMD recommends export HIP_VISIBLE_DEVICES=0 within each VF context to restrict GPU visibility.

04

NVIDIA VGPU: SOFTWARE VIRTUALIZATION OVERHEAD

NVIDIA vGPU (formerly GRID) uses a hypervisor-resident virtual GPU manager that intercepts GPU commands from guest VMs and schedules them on the physical GPU. Unlike MIG's hardware partitioning, vGPU is a software abstraction layer that introduces 8-25% performance overhead depending on workload. Compute-heavy workloads (GEMM, convolutions) experience 8-12% overhead because the vGPU manager's command interception adds 10-25 microseconds per kernel launch. Memory-bandwidth-bound workloads (attention, embedding lookups) experience 15-25% overhead because the vGPU manager mediates all GPU memory accesses through the hypervisor's IOMMU.

vGPU supports frame buffer reservation, which guarantees a minimum GPU memory allocation per VM. The key operational advantage is live migration: vGPU supports VM live migration between physical hosts with vGPU-licensed NVIDIA GPUs, which MIG and SR-IOV do not. For enterprise VDI workloads requiring GPU-accelerated desktops, vGPU remains the standard. For AI workloads, the 8-25% overhead and per-VM licensing cost ($2,500-3,500 per GPU per year) make it inferior to MIG for production inference and training.

Virtualization MethodCompute OverheadMemory OverheadKernel Launch OverheadMemory BW LossLive Migration
MIG (hardware slice)0.2-0.4%0.1-0.3%0.5-1.0 us0.1-0.3%No
AMD SR-IOV1.5-3.8%2.0-4.0%1.5-3.0 us2-4%No
NVIDIA vGPU (GRID)8-12%15-25%10-25 us15-25%Yes
K8s Time-SlicingVariable0% (shared memory)0% (direct pass-through)0%No
GPU Passthrough (PCIe)0%0%0%0%No
05

KUBERNETES GPU SHARING: TIME-SLICING AND MIG IN PRACTICE

For Kubernetes GPU clusters, the most common virtualization approaches are MIG (for NVIDIA GPUs), SR-IOV (for AMD GPUs), and time-slicing (for any GPU without hardware partitioning). Time-slicing via the NVIDIA GPU Operator's TimeSlicing config allocates GPU time to pods on a best-effort basis. The default scheduling quantum is 50ms: each pod gets exclusive GPU access for 50ms before the scheduler preempts and context-switches to the next pod. The overhead of time-slicing is highly workload-dependent: compute-bound workloads see 3-5% overhead per additional sharing pod.

MIG on Kubernetes is managed through the nvidia.com/gpu resource with the nvidia.com/mig-profile annotation. A pod requesting nvidia.com/gpu: 1 with annotation nvidia.com/mig-profile: 1g.11gb receives a dedicated MIG slice. Our production benchmark on 8x H100 with 56 MIG slices running independent Llama 8B inference workloads: each slice serves 22 tok/s, aggregate throughput 1,232 tok/s, versus 4,200 tok/s on 8 full H100 GPUs.

06

DEPLOYMENT RECOMMENDATIONS BY WORKLOAD TYPE

The optimal GPU virtualization strategy depends on workload characteristics. For production LLM inference serving with strict latency SLAs, use MIG or SR-IOV with dedicated GPU slices per model replica. The hardware isolation guarantees no noisy-neighbor effects. For batch inference and fine-tuning with relaxed latency requirements, time-slicing with the Kubernetes GPU Operator is cost-effective, achieving 60-75% utilization on H100 clusters. For VDI and GPU-accelerated desktop workloads, NVIDIA vGPU with GRID licensing is the only option that supports live migration.

On ClusterBid, GPU instances are tagged with their virtualization capabilities via the --gpu-partitioning filter with values mig, sriov, vgpu, timeslice, or passthrough. For cost optimization, compare MIG-partitioned H100 instances ($2.50-3.50/hour per 1/7 slice) against full H100 instances ($3.50-5.00/hour) with software time-slicing. At 70% utilization, MIG slices provide 35% better cost efficiency for models fitting within 11 GB of GPU memory.

Filed under
GPU VirtualizationNVIDIA vGPUMIG PartitioningSR-IOV GPUGPU Performance OverheadGPU IsolationH100 MIG