All essays
MarketMARKET REPORTFEB 2026

GPU Fractional Compute: How MIG and GPU Partitioning Slash Costs for Small Model Inference and Development

Analysis of NVIDIA MIG, AMD MxGPU, and software-level GPU partitioning for inference and development workloads. Real cost savings data for small model serving, CI/CD test suites, and dev environments.

01

The Full-GPU Waste Problem

Most GPU deployments allocate full GPUs to every workload. For serving Llama 3.2 3B or Qwen 2.5 7B, a full H100 at $3.10/GPU/hr is wildly overprovisioned. These models use 10-30 GB of VRAM for inference with batch size 1. An H100 has 80 GB. The remaining 50-70 GB sits idle, consuming power and generating heat. At scale, this waste compounds: a 1,000-GPU cluster running 7B parameter inference wastes roughly $1.5M per year in unused compute capacity.

The development gap is even wider. GPU CI/CD pipelines that run model validation, unit tests with GPU kernels, and small batch training runs for hyperparameter sweeps rarely need full GPU resources. A typical CI job on an H100 uses 4-8 GB of VRAM and 15-25% of SM compute capacity. Without partitioning, each CI job requires an exclusive GPU, creating contention and queue times for the team.

02

NVIDIA MIG: Hardware-Level Partitioning

Multi-Instance GPU (MIG) on H100 and B300 GPUs partitions the GPU at the hardware level into up to 7 instances on an H100 80 GB (or up to 7 on B300 288 GB depending on profile). Each instance gets dedicated SM slices, dedicated HBM slices, and a separate L2 cache partition. There is no resource contention between instances because the memory controllers enforce physical partitioning. GPU instances each appear as distinct CUDA devices to the OS.

MIG supports profiles of 1g.10gb (1 SM slice, 10 GB), 2g.20gb, 3g.40gb, 4g.40gb, and 7g.80gb (full GPU). For inference serving, the 1g.10gb profile runs a quantized Llama 3.2 3B (4-bit AWQ, 2.1 GB VRAM) with 20-30 ms per token latency. A single H100 can serve 7 independent inference endpoints at $0.44/GPU/hr per endpoint versus $3.10/GPU/hr for a dedicated GPU. That is roughly 86% cost reduction per inference deployment.

MIG ProfileSM SlicesHBMMax Instances (H100 80GB)Target Workload
1g.10gb110 GB73B-7B quantized inference
2g.20gb220 GB47B-13B inference, dev containers
3g.40gb340 GB213B-30B inference, small training
4g.40gb440 GB1 (plus others)30B-70B inference, LoRA training
7g.80gb780 GB1Full GPU, large training runs
03

Software Partitioning: Time-Slicing and MPS

NVIDIA MPS (Multi-Process Service) provides software-level GPU sharing by merging multiple CUDA streams into a single submission context. Each process gets proportional SM time but shares the full HBM pool. MPS is ideal for bursty development workloads where 5-10 developers run short training jobs concurrently. Unlike MIG, there is no HBM isolation, so one process can OOM the entire GPU. MPS achieves roughly 90% spatial utilization versus MIG's 95%+.

Time-slicing through Kubernetes device plugin scheduling is the simplest approach: multiple pods share a GPU sequentially, each getting exclusive access for a time quantum. This works well for interactive development (Jupyter notebooks, VS Code remote) where latency tolerance is high. However, time-slicing causes severe latency spikes for inference serving (50-500 ms tail latency increases under contention). For production inference, MIG or MPS with QoS guarantees is required.

04

Cost Model: Partitioned vs Full GPU

The economics of GPU partitioning are straightforward. A development team of 10 engineers running model experimentation, fine-tuning, and CI/CD on a single H100 partitioned into 4x 2g.20gb MIG instances at $0.78/hr each charges the business $0.78/hr total (4 instances active). A dedicated H100 per engineer at $3.10/hr each would cost $31.00/hr. The monthly cost difference: approximately $560 versus $22,300 at 24/7 utilization.

For inference, the math depends on request volume. Serving 100 concurrent requests for a 7B model requires roughly 4 H100 GPUs in dedicated mode ($12.40/hr). With MIG 2g.20gb profiles running vLLM with continuous batching, 8 instances on 2 H100s handle the same throughput at $6.20/hr. The tradeoff is maximum batch size per instance: MIG profiles limit per-instance throughput by capping SM count, so very high-throughput endpoints may need multiple instances load-balanced.

WorkloadFull GPU Cost/moPartitioned Cost/moSavingsBest Method
Dev team (10 engineers)$22,300$56097%MIG 2g.20gb
7B inference (100 req/s)$8,928$4,46450%MIG 2g.20gb x8
CI/CD GPU test suite$2,232$28087%MPS time-shared
LoRA fine-tuning (13B)$2,232$1,11650%MIG 3g.40gb
Batch data preprocessing$2,232$56075%MPS concurrent
05

MIG Limitations and Caveats

MIG has critical constraints. NCCL inter-instance communication is not supported: MIG instances cannot participate in collective operations (all-reduce, all-gather) with each other or with other GPUs. This makes multi-GPU training impossible within MIG instances. For training workflows that require data parallelism or model parallelism, MIG is unsuitable. Teams requiring both inference and training on the same hardware should mix MIG for inference with full-GPU partitions for training.

MIG is also limited to NVIDIA Ampere and newer GPUs (A100, H100, B300) and requires compatible drivers (R525+), CUDA 11.7+, and a GPU with MIG-enabled firmware. Not all GPU SKUs support MIG. H100 PCIe supports MIG with up to 7 instances. H100 SXM supports 7 instances. B300 supports up to 7 instances per GPU, but MIG on B300 requires CUDA 13.0+. MIG also disables NVLink peer-to-peer access between instances, which impacts GPU-to-GPU communication for distributed inference.

06

Implementation Strategy

For inference deployments, use MIG with Kubernetes and the NVIDIA MIG Manager. Deploy the MIG partition daemon as a DaemonSet that creates and destroys partitions based on Pod GPU requests. Use the `nvidia.com/mig-profile` resource label to request specific profiles. Route inference requests to endpoints via Istio or Kong with MIG instance awareness. Monitor per-instance metrics with DCGM (available per MIG slice via `nvidia-smi mig -i GPUID -gi INSTANCEID`).

For development environments, use MPS with resource limits in Docker. Launch containers with `--gpus '"device=0"'` and set `NVIDIA_MPS_CONTROL_DIR` environment variables. Set per-process memory limits via `CUDA_VISIBLE_DEVICES` and `nvidia-smi --gpu-reset`. For CI/CD, prefer MPS over MIG because MIG requires pre-allocating profiles while MPS dynamically shares resources. A cluster running 8-16 MPS-shared GPUs typically serves 50-100 developer CI jobs daily without contention.

Filed under
MIGGPU partitioningfractional computeinference costH100multi-tenant GPUdevelopment environmentsGPU utilization