The Full-GPU Waste Problem
Most GPU deployments allocate full GPUs to every workload. For serving Llama 3.2 3B or Qwen 2.5 7B, a full H100 at $3.10/GPU/hr is wildly overprovisioned. These models use 10-30 GB of VRAM for inference with batch size 1. An H100 has 80 GB. The remaining 50-70 GB sits idle, consuming power and generating heat. At scale, this waste compounds: a 1,000-GPU cluster running 7B parameter inference wastes roughly $1.5M per year in unused compute capacity.
The development gap is even wider. GPU CI/CD pipelines that run model validation, unit tests with GPU kernels, and small batch training runs for hyperparameter sweeps rarely need full GPU resources. A typical CI job on an H100 uses 4-8 GB of VRAM and 15-25% of SM compute capacity. Without partitioning, each CI job requires an exclusive GPU, creating contention and queue times for the team.
NVIDIA MIG: Hardware-Level Partitioning
Multi-Instance GPU (MIG) on H100 and B300 GPUs partitions the GPU at the hardware level into up to 7 instances on an H100 80 GB (or up to 7 on B300 288 GB depending on profile). Each instance gets dedicated SM slices, dedicated HBM slices, and a separate L2 cache partition. There is no resource contention between instances because the memory controllers enforce physical partitioning. GPU instances each appear as distinct CUDA devices to the OS.
MIG supports profiles of 1g.10gb (1 SM slice, 10 GB), 2g.20gb, 3g.40gb, 4g.40gb, and 7g.80gb (full GPU). For inference serving, the 1g.10gb profile runs a quantized Llama 3.2 3B (4-bit AWQ, 2.1 GB VRAM) with 20-30 ms per token latency. A single H100 can serve 7 independent inference endpoints at $0.44/GPU/hr per endpoint versus $3.10/GPU/hr for a dedicated GPU. That is roughly 86% cost reduction per inference deployment.
| MIG Profile | SM Slices | HBM | Max Instances (H100 80GB) | Target Workload |
|---|---|---|---|---|
| 1g.10gb | 1 | 10 GB | 7 | 3B-7B quantized inference |
| 2g.20gb | 2 | 20 GB | 4 | 7B-13B inference, dev containers |
| 3g.40gb | 3 | 40 GB | 2 | 13B-30B inference, small training |
| 4g.40gb | 4 | 40 GB | 1 (plus others) | 30B-70B inference, LoRA training |
| 7g.80gb | 7 | 80 GB | 1 | Full GPU, large training runs |
Software Partitioning: Time-Slicing and MPS
NVIDIA MPS (Multi-Process Service) provides software-level GPU sharing by merging multiple CUDA streams into a single submission context. Each process gets proportional SM time but shares the full HBM pool. MPS is ideal for bursty development workloads where 5-10 developers run short training jobs concurrently. Unlike MIG, there is no HBM isolation, so one process can OOM the entire GPU. MPS achieves roughly 90% spatial utilization versus MIG's 95%+.
Time-slicing through Kubernetes device plugin scheduling is the simplest approach: multiple pods share a GPU sequentially, each getting exclusive access for a time quantum. This works well for interactive development (Jupyter notebooks, VS Code remote) where latency tolerance is high. However, time-slicing causes severe latency spikes for inference serving (50-500 ms tail latency increases under contention). For production inference, MIG or MPS with QoS guarantees is required.
Cost Model: Partitioned vs Full GPU
The economics of GPU partitioning are straightforward. A development team of 10 engineers running model experimentation, fine-tuning, and CI/CD on a single H100 partitioned into 4x 2g.20gb MIG instances at $0.78/hr each charges the business $0.78/hr total (4 instances active). A dedicated H100 per engineer at $3.10/hr each would cost $31.00/hr. The monthly cost difference: approximately $560 versus $22,300 at 24/7 utilization.
For inference, the math depends on request volume. Serving 100 concurrent requests for a 7B model requires roughly 4 H100 GPUs in dedicated mode ($12.40/hr). With MIG 2g.20gb profiles running vLLM with continuous batching, 8 instances on 2 H100s handle the same throughput at $6.20/hr. The tradeoff is maximum batch size per instance: MIG profiles limit per-instance throughput by capping SM count, so very high-throughput endpoints may need multiple instances load-balanced.
| Workload | Full GPU Cost/mo | Partitioned Cost/mo | Savings | Best Method |
|---|---|---|---|---|
| Dev team (10 engineers) | $22,300 | $560 | 97% | MIG 2g.20gb |
| 7B inference (100 req/s) | $8,928 | $4,464 | 50% | MIG 2g.20gb x8 |
| CI/CD GPU test suite | $2,232 | $280 | 87% | MPS time-shared |
| LoRA fine-tuning (13B) | $2,232 | $1,116 | 50% | MIG 3g.40gb |
| Batch data preprocessing | $2,232 | $560 | 75% | MPS concurrent |
MIG Limitations and Caveats
MIG has critical constraints. NCCL inter-instance communication is not supported: MIG instances cannot participate in collective operations (all-reduce, all-gather) with each other or with other GPUs. This makes multi-GPU training impossible within MIG instances. For training workflows that require data parallelism or model parallelism, MIG is unsuitable. Teams requiring both inference and training on the same hardware should mix MIG for inference with full-GPU partitions for training.
MIG is also limited to NVIDIA Ampere and newer GPUs (A100, H100, B300) and requires compatible drivers (R525+), CUDA 11.7+, and a GPU with MIG-enabled firmware. Not all GPU SKUs support MIG. H100 PCIe supports MIG with up to 7 instances. H100 SXM supports 7 instances. B300 supports up to 7 instances per GPU, but MIG on B300 requires CUDA 13.0+. MIG also disables NVLink peer-to-peer access between instances, which impacts GPU-to-GPU communication for distributed inference.
Implementation Strategy
For inference deployments, use MIG with Kubernetes and the NVIDIA MIG Manager. Deploy the MIG partition daemon as a DaemonSet that creates and destroys partitions based on Pod GPU requests. Use the `nvidia.com/mig-profile` resource label to request specific profiles. Route inference requests to endpoints via Istio or Kong with MIG instance awareness. Monitor per-instance metrics with DCGM (available per MIG slice via `nvidia-smi mig -i GPUID -gi INSTANCEID`).
For development environments, use MPS with resource limits in Docker. Launch containers with `--gpus '"device=0"'` and set `NVIDIA_MPS_CONTROL_DIR` environment variables. Set per-process memory limits via `CUDA_VISIBLE_DEVICES` and `nvidia-smi --gpu-reset`. For CI/CD, prefer MPS over MIG because MIG requires pre-allocating profiles while MPS dynamically shares resources. A cluster running 8-16 MPS-shared GPUs typically serves 50-100 developer CI jobs daily without contention.
