Why GPU Sharing Matters
GPU utilisation in AI clusters is often surprisingly low. Training workloads achieve 40-55% MFU, meaning 45-60% of GPU compute capacity is idle during training. Inference workloads achieve 15-35% utilisation because memory bandwidth constraints limit compute utilisation. The gap between GPU purchase cost and GPU utilisation represents a massive efficiency opportunity.
GPU sharing - running multiple workloads on the same physical GPU - improves utilisation by filling idle cycles. The challenge is maintaining performance isolation, ensuring that one workload does not degrade another's performance, and managing GPU memory allocation between competing workloads. At mid-2026, several GPU sharing technologies address this challenge at different levels of the stack.
This post covers the GPU sharing landscape: NVIDIA vGPU, MIG partitioning, SR-IOV virtualization, GPU time-slicing, and software-level sharing through CUDA MPS. Each approach offers different isolation guarantees, performance characteristics, and use case suitability.
MIG: Hardware-Enforced GPU Partitioning
Multi-Instance GPU (MIG) is NVIDIA's hardware-enforced partitioning technology, available on A100, H100, and B200. MIG partitions the GPU into up to 7 fully isolated instances, each with dedicated memory, cache, and compute units. The isolation is hardware-enforced - there is no shared pathway between MIG instances, making MIG the strongest GPU sharing technology for multi-tenant environments.
MIG partitioning profiles for H100: 1g.10gb (1 SM, 10 GB memory, 7 instances per GPU), 2g.20gb (2 SMs, 20 GB, 3 instances), 3g.40gb (3 SMs, 40 GB, 2 instances), and 7g.80gb (full GPU). For B200 (192 GB): up to 7 instances with 24 GB minimum per instance.
MIG is ideal for multi-tenant environments with strong isolation requirements: SaaS platforms serving multiple customers, regulated workloads requiring tenant separation, and production plus staging on the same GPU. The trade-off is reduced flexibility - MIG partitions are static and cannot be resized without GPU reset.
vGPU: Virtual GPU for Virtualized Environments
NVIDIA vGPU enables GPU sharing in virtualized environments (VMware vSphere, KVM with QEMU). vGPU allocates a portion of GPU memory and compute to each virtual machine, with time-sliced GPU scheduling. Unlike MIG, vGPU does not provide hardware-enforced memory isolation - it relies on the hypervisor and GPU driver for memory protection.
vGPU is available in two modes: time-sliced (each VM gets a fraction of GPU time, all VMs share the full GPU memory) and vGPU with MIG (MIG partitions exposed to VMs, providing hardware isolation). The vGPU-with-MIG mode is the preferred configuration for multi-tenant virtualized GPU environments.
At mid-2026, vGPU is primarily used for virtual desktop infrastructure with AI capabilities. vGPU's time-sliced scheduling causes higher latency variance than MIG, making it unsuitable for latency-sensitive inference serving. For batch workloads and interactive desktop use cases, vGPU provides adequate performance at lower cost.
GPU Time-Slicing: Simple Sharing with Limitations
GPU time-slicing is the simplest GPU sharing approach: the GPU scheduler divides GPU time between multiple processes, with each process getting exclusive GPU access for a time quantum (typically 10-100ms). Time-slicing is supported by the Kubernetes NVIDIA device plugin, NVIDIA MPS (Multi-Process Service), and CUDA streams.
Time-slicing is not true isolation. The GPU memory is shared between all time-sliced processes, meaning one process can cause OOM errors for others. Compute isolation is also limited - a compute-heavy workload can consume GPU cycles and degrade other workloads' performance. Time-slicing is best suited for co-operative workloads from the same tenant.
Kubernetes time-slicing configuration: the NVIDIA device plugin can be configured to expose multiple GPU time-slice resources from a single GPU (e.g., expose a physical GPU as 5 virtual GPUs, each getting 20% of GPU time). This enables oversubscribing GPUs for development environments, CI/CD pipelines, and batch processing where throughput consistency is not required.
SR-IOV: Direct GPU Pass-Through for Virtualization
Single Root I/O Virtualization (SR-IOV) for GPUs enables direct pass-through of GPU virtual functions to virtual machines, bypassing the hypervisor's GPU scheduler. Each VM gets near-native GPU performance because the GPU is accessed directly, not through a software virtualization layer.
SR-IOV is less flexible than vGPU or MIG: it does not support GPU memory partitioning or oversubscription. Each GPU function is assigned entirely to one VM. SR-IOV is best suited for workloads that need near-native GPU performance in a virtualized environment, such as GPU-accelerated databases or real-time inference at the edge.
At mid-2026, SR-IOV for GPUs is not widely deployed in cloud AI infrastructure. MIG and vGPU provide better resource utilisation for the multi-tenant GPU cluster model. SR-IOV has found its niche in edge computing and telecommunications, where workloads are stable and the overhead of software virtualization is unacceptable.
CUDA MPS: Software-Level GPU Sharing
The CUDA Multi-Process Service (MPS) provides software-level GPU sharing for multiple CUDA processes. MPS uses a client-server architecture: a single MPS server process manages GPU access for multiple client processes, providing cooperative scheduling of GPU compute resources.
MPS improves GPU utilisation by allowing concurrent kernel execution from multiple processes on the same GPU. It is particularly effective for inference workloads where multiple model replicas share a GPU. MPS does not provide memory isolation - all client processes share the GPU memory space - so it is not suitable for untrusted multi-tenant environments.
In practice, MPS is used to increase GPU utilisation for single-tenant serving infrastructure. A typical deployment runs 4-8 model replicas per GPU through MPS, achieving 60-80% GPU utilisation compared to 20-30% for single-replica deployments on the same hardware.
Choosing the Right GPU Sharing Strategy
The GPU sharing strategy depends on isolation requirements, workload characteristics, and management overhead tolerance. The decision matrix: MIG (strongest isolation, static partitioning, ideal for regulated multi-tenant workloads). vGPU (good isolation in virtualized environments, flexible allocation, good for VDI). Time-slicing (weakest isolation, highest flexibility, good for dev/staging and non-critical workloads). CUDA MPS (shared memory space, cooperative scheduling, good for single-tenant multi-replica inference).
Hybrid approaches are common in production. A typical 256-GPU cluster might use: MIG for 40% of GPUs (multi-tenant inference), single-GPU for 30% (large model training), MPS for 20% (single-tenant multi-replica inference), and time-slicing for 10% (CI/CD and experimentation).
The trend at mid-2026 is toward more flexible GPU sharing. NVIDIA's software-defined GPU orchestration (introduced with Blackwell) enables dynamic GPU partitioning without hardware reset, combining the isolation of MIG with the flexibility of time-slicing. This software-defined approach is expected to become the dominant GPU sharing model by 2027.
