MIG TECHNOLOGY EXPLAINED
Multi-Instance GPU (MIG) partitions a single NVIDIA H100 or H200 into up to 7 independent GPU instances. Each MIG slice has dedicated compute units, memory bandwidth, and L2 cache with full fault isolation. The GPU's 80 GB HBM3 is divided into profiles: 1g.10gb (1 compute unit, 10 GB), 2g.20gb, 3g.40gb, 4g.40gb, and 7g.80gb (full GPU).
MIG is supported on H100 SXM and PCIe, H200, and the upcoming B200 (with updated MIG 2.0). AMD's equivalent technology, called GPU Partitioning on MI300X, provides similar isolation with 2x 96 GB or 4x 48 GB partitions. The key constraint: all MIG slices on a GPU must use the same model due to shared memory controller configuration.
MIG PARTITION PROFILES FOR INFERENCE
For 7B models at FP8 (14 GB weights + 4-8 GB KV cache), a 3g.40gb MIG slice provides comfortable headroom. Cost: ~$0.80-1.20/hr per slice versus $2.50/hr for full H100. A single H100 with 2x 3g.40gb and 1x 1g.10gb slices serves 3 inference endpoints simultaneously at 55-65% of full GPU throughput.
For 13B models at FP8 (26 GB weights + 6-10 GB KV cache), a 4g.40gb slice is required. Two 4g.40gb slices per H100 = 2 inference endpoints. For 34B models at FP8 (68 GB), only a full 7g.80gb slice works unless using INT4 quantization (34 GB weights), which enables 4g.40gb deployment with batch size up to 16.
| Model Size | Precision | VRAM Req | MIG Profile | Slices per H100 | $/slice/hr |
|---|---|---|---|---|---|
| 7B | FP8 | ~20 GB | 3g.40gb | 2 | $0.80-1.20 |
| 7B | INT4 | ~10 GB | 1g.10gb | 7 | $0.30-0.50 |
| 13B | FP8 | ~34 GB | 4g.40gb | 2 | $1.20-1.60 |
| 13B | INT4 | ~18 GB | 3g.40gb | 2 | $0.80-1.20 |
| 34B | FP8 | ~74 GB | 7g.80gb | 1 | $2.50 |
| 34B | INT4 | ~40 GB | 4g.40gb | 2 | $1.20-1.60 |
MIG INFERENCE DEPLOYMENT PATTERNS
Each MIG slice runs its own vLLM or TGI instance with independent port, rate limiting, and model configuration. Kubernetes with NVIDIA MIG device plugin manages MIG slices as GPU resources, enabling pod-per-slice scheduling. Monitoring per-slice GPU utilization is critical: MIG hides per-slice metrics behind nvidia-smi MIG mode.
The optimal deployment pattern groups models of the same size on a single GPU. For example, two 13B endpoints on two 4g.40gb slices of one H100, sharing the same model weights via NCCL. This weight-sharing pattern reduces total VRAM consumption by 20-30% compared to independent model loads.
COST SAVINGS ANALYSIS
A team running eight 7B inference endpoints on H100 currently deploys 8 full GPUs at $20/hr ($2.50 x 8). With MIG 3g.40gb slices running 2 per H100, the same workload requires 4x H100 at $10/hr. Annual savings at 50% utilization: $43,800 per year. For 13B endpoints (8 endpoints, 2 per H100 via 4g.40gb), savings reach $21,900 annually.
The savings compound for teams running mixed small model fleets. A deployment of 6x 7B, 4x 13B, and 2x 34B endpoints currently requiring 12 H100s at $30/hr can fit on 7 H100s at $17.50/hr with MIG mapping, a 42% cost reduction. The MIG configuration requires upfront planning but pays for itself within the first month.
WORKLOAD-TO-MIG MAPPING STRATEGY
Not all workloads are MIG-compatible. Batch inference and low-concurrency serving work well. High-throughput serving with batch sizes above 64 prefers full GPU because MIG's memory bandwidth partitioning becomes the bottleneck. MIG slices suffer 15-25% per-token throughput penalty at high batch sizes due to shared HBM bandwidth contention.
Recommended mapping: latency-sensitive small models (7B, batch < 32) on 3g.40gb slices, throughput-oriented medium models (13B, batch < 64) on 4g.40gb slices, and high-throughput large models (34B+) on full GPUs. Batch processing of small models (7B, batch > 64) also prefers full GPU despite VRAM headroom.
MIG IMPLEMENTATION GUIDE
Enable MIG on H100 via nvidia-smi: `nvidia-smi mig -cgi 19,19,14 -C` creates one 4g.40gb and two 3g.40gb slices. The MIG configuration is persistent across reboots. Kubernetes setup requires NVIDIA MIG device plugin v0.9+ and NVIDIA GPU Operator 24.9+. For neoclouds, RunPod launched MIG support in May 2026; Lambda plans Q3 2026 support.
Key gotchas: MIG CUDA context switching overhead adds 200-500ms when reconfiguring slices. All slices on a GPU must use the same NCCL domain, preventing inter-slice communication. MIG slices cannot be dynamically resized; reprovisioning requires stopping all slices. Plan MIG configurations as fixed deployments, not dynamic pools.
