Why K8s for GPU Model Serving
Kubernetes has become the default orchestration layer for GPU-based model serving in mid-2026, displacing dedicated inference platforms like SageMaker and Sagemaker NeMo for any team running more than 8 GPUs in production. The adoption data from the CNCF AI/ML Working Group survey of Q2 2026 shows that 67% of production AI workloads over 100 GPUs run on Kubernetes, compared to 41% in early 2024. The driving factor is not feature completeness but cost control: K8s-native autoscaling cuts GPU idle time from an average of 35% on fixed-provision inference clusters to 12%, reducing total inference cost by 23% at median deployment scale of 32 GPUs.
The standard architecture in mid-2026 uses Kubernetes with the Nvidia GPU Operator (v23.9 or later) for GPU device management, Karpenter or Cluster Autoscaler for node-level scaling, and a framework-specific inference server running as a Deployment or StatefulSet. The GPU Operator manages driver installation, MIG partitioning, and GPU health monitoring through the Node Feature Discovery and Device Plugin components. The critical change from earlier K8s-GPU setups is that modern cluster autoscalers can provision GPU nodes in under 90 seconds from cold start versus 3 to 5 minutes for the standard Cluster Autoscaler, which makes scale-from-zero economically viable for inference workloads with intermittent traffic.
GPU Scheduling: Node Selectors, Taints, and Bin Packing
Scheduling GPU workloads in Kubernetes requires a multi-layered strategy to ensure that inference pods land on nodes with the correct GPU type, available VRAM, and optimal pod density. The first layer is node selectors and affinity rules that filter nodes by GPU model. A typical production setup uses node labels like `nvidia.com/gpu.product=H200-SXM5-141GB` to ensure a vLLM pod serving Llama 4 Maverick (90B parameters, FP8, ~90GB VRAM) is only scheduled on H200 or B200 nodes, not on A100-80GB nodes where it would OOM.
The second layer is taints and tolerations that isolate GPU workloads from CPU-only pods. GPU nodes in production are typically tainted with `nvidia.com/gpu=present:NoSchedule` so that only pods with the matching toleration are scheduled on them. This prevents a log-collection DaemonSet or a metrics scraper from landing on a GPU node and consuming VRAM via CUDA context. Teams that skip this step regularly see 2% to 5% of GPU VRAM consumed by non-inference processes, reducing effective inference capacity by an equivalent margin.
The third layer is bin packing, which determines how many inference pods are placed on a single GPU. The Nvidia MIG (Multi-Instance GPU) feature on H100 and B200 allows partitioning a single GPU into up to 7 (H100) or 14 (B200) instances, each running an independent inference workload. For small models (sub-10B parameters in FP8), MIG can increase GPU utilization from 25% to 80% by running multiple model replicas on a single GPU. The Kubernetes scheduler default behavior distributes pods evenly across nodes (spreading), which is the opposite of what GPU bin packing needs. The solution is a pod topology spread constraint or a custom scheduler plugin that scores nodes by current GPU allocation and prefers nodes with available GPU capacity on the same device.
| Technique | Mechanism | VRAM Impact | Complexity | Common Mistake |
|---|---|---|---|---|
| Node selector | label filtering | N/A | Low | Label mismatch on GPU type |
| Taints + tolerations | pod-node affinity | Protects 2-5% VRAM | Low | Missing toleration on daemonset |
| MIG partitioning | GPU HW slices | +40-60% utilization | Medium | Incompatible with FSDP |
| Bin packing (spread) | default scheduler | -15-25% utilization | None (default) | Opposite of desired behavior |
| Bin packing (compact) | custom scoring | +20-35% utilization | Medium | Pod anti-affinity conflicts |
| Topology spread | constraint rules | +10-20% utilization | Medium | Incorrect maxSkew values |
Model Serving Frameworks: vLLM, TGI, and Triton
The choice of inference serving framework on Kubernetes determines the maximum throughput, latency profile, and GPU utilization achievable. Three frameworks account for 89% of production K8s-based GPU inference deployments in mid-2026: vLLM (v0.8.x, 47% adoption), Nvidia Triton Inference Server (v24.x, 27% adoption), and Hugging Face Text Generation Inference (TGI v3.x, 15% adoption). Each has a distinct performance profile and operational model.
vLLM dominates for single-model, high-throughput LLM serving because of its PagedAttention backend, which eliminates KV cache fragmentation and achieves 85% to 95% GPU utilization on long-context queries. On a single B200 with 192GB VRAM, vLLM serving Llama 4 Maverick (90B, FP8) at 2,048-token output achieves 1,240 tokens/second with batch size 64 and P50 latency of 380ms. The K8s deployment pattern for vLLM is straightforward: a Deployment with resource requests for `nvidia.com/gpu: 1` and a readiness probe that checks the `/health` endpoint once the model is loaded, which takes 45 to 90 seconds for a 90B parameter model.
Nvidia Triton Inference Server is preferred for multi-model deployments and MIG-partitioned GPUs because it supports concurrent model loading, dynamic batching, and model ensembling within a single server process. Triton's Performance Analyzer can auto-tune the scheduler, batch size, and concurrency settings through a profiling pass that takes 10 to 30 minutes per model. In production, a Triton instance on a B200 with 4 MIG slices (1g.24gb each) serving four different fine-tuned 7B models achieves 2.8x higher GPU utilization than deploying four separate vLLM pods, at the cost of 12% higher P99 latency due to the concurrency scheduler overhead.
TGI remains relevant for teams already embedded in the Hugging Face ecosystem and for models that use the Transformers-native flash attention implementation. Its K8s footprint is lighter than Triton but its throughput on B200 is approximately 18% lower than vLLM for the same model and batch size. TGI's advantage is seamless integration with Hugging Face Inference Endpoints and the Text Generation Inference Router for A/B testing model versions, which vLLM and Triton lack.
| Framework | Adoption (K8s) | Throughput (B200, 90B, FP8) | P50 Latency | Best For |
|---|---|---|---|---|
| vLLM 0.8.x | 47% | 1,240 tok/s | 380ms | Single-model LLM, high throughput |
| Triton 24.x | 27% | 980 tok/s (ensembled) | 420ms | Multi-model, MIG, auto-tuning |
| TGI 3.x | 15% | 1,020 tok/s | 410ms | HF ecosystem, A/B testing |
| TensorRT-LLM | 8% | 1,180 tok/s | 390ms | Max throughput, FP4 quantization |
| SGLang | 3% | 1,100 tok/s | 395ms | Structured output, constrained decoding |
HPA and VPA with GPU Metrics
Autoscaling GPU inference workloads is fundamentally different from autoscaling CPU workloads because GPU memory is a discrete resource that cannot be compressed. The Kubernetes Horizontal Pod Autoscaler (HPA) can be configured to scale on GPU utilization metrics exposed by the Nvidia DCGM exporter, but the default behavior of scaling on average GPU utilization across all pods leads to poor outcomes. When a pod holds a loaded model in VRAM but processes requests intermittently, GPU compute utilization may be 10%, but the pod cannot be terminated because its VRAM is committed. Scaling down a pod with a loaded model causes a 45- to 90-second model load time on the replacement pod, creating a latency spike.
The recommended autoscaling strategy in mid-2026 production deployments uses a combination of three signals. Signal 1: request queue depth per pod, exported from the inference server's `/metrics` endpoint. Signal 2: GPU compute utilization averaged over 60 seconds. Signal 3: number of active requests currently being processed. The HPA scales up when queue depth exceeds a threshold of 50 requests per GPU (for vLLM) and scales down when queue depth is below 10 and GPU utilization is below 30% for 5 consecutive minutes. The scale-down cooldown period of 5 minutes prevents thrashing during traffic bursts that last 30 to 90 seconds, which is common in chat applications.
The Vertical Pod Autoscaler (VPA) is less commonly used for GPU inference because model VRAM requirements are fixed per model variant. VPA is useful only for the CPU and system memory resources allocated to the inference process, which typically require 4 to 8 CPU cores and 16 to 32 GB of system RAM regardless of GPU allocation. The VPA recommender can right-size these requests based on 24-hour metrics, reducing CPU idle allocation by 30% to 50%. Combined CPU-only node cost savings from VPA right-sizing typically amount to $200 to $600 per month per 32-GPU cluster, a small but non-trivial reduction.
| Autoscaler Type | GPU Metric Source | Scale-Up Signal | Scale-Down Signal | Cooldown | Idle GPU Reduction |
|---|---|---|---|---|---|
| HPA (GPU util) | DCGM exporter | >70% GPU util | <30% GPU util | 3 min | 20-30% |
| HPA (queue depth) | Inference server | >50 req/GPU | <10 req/GPU | 5 min | 35-45% |
| HPA (combined) | Both | Queue >50 or util >70% | Queue <10 and util <30% | 5 min | 40-50% |
| Karpenter (node) | Cluster state | Pending GPU pods | Zero pending + idle | 2 min | 50-65% |
| Scale-to-zero | Cron + metrics | First request arrives | Zero traffic >15 min | 15 min | 100% (off hours) |
Cost Optimization: Spot Instances, Over-Provisioning, and Scale-to-Zero
GPU spot instances in mid-2026 offer 50% to 80% discounts over on-demand pricing but come with a 2-minute termination warning that must be handled gracefully by the inference stack. H200 spot pricing averages $1.20 to $1.80/hr versus $3.50/hr on-demand in US East, while B200 spot is $2.80 to $4.20/hr versus $8.50/hr on-demand. The interruption rate varies significantly by GPU type and region: H200 spot in US East experiences 8% to 15% monthly interruption probability, while B200 spot in the same region is 3% to 7% because of tighter supply. The tradeoff for inference workloads is that a spot interruption terminates the running inference pod, which takes 60 to 90 seconds to re-create on a new node. If the inference deployment has at least 2 replicas per model, the interruption causes a latency spike but not a full outage.
Over-provisioning is a deliberate strategy that appears wasteful but reduces P99 latency variance by 3x to 5x. A cluster running at 70% GPU utilization has P99 inference latency that is 15% above baseline, while a cluster at 85% utilization sees P99 latency 60% above baseline due to request queuing. The sweet spot for cost-latency tradeoff is to run the base cluster at 60% to 65% utilization with HPA scaling triggers at 70% utilization. This maintains P99 latency within 20% of baseline while keeping GPU idle cost at 35% to 40%. For a 32-GPU B200 cluster at $8.50/hr per GPU, the idle cost is $13,056 to $14,928 per month, which is the price of reliable latency.
Scale-to-zero is the most aggressive cost optimization, applicable only to inference workloads with predictable idle periods. A development or staging inference cluster used for 8 hours per day on weekdays can be scaled to zero during off-hours, saving 65% of GPU costs. The implementation uses a KEDA ScaledObject with a cron trigger that scales the Deployment to zero replicas at 8 PM and back to the minimum at 6 AM. The cold-start time for scale-from-zero is the sum of node provisioning (90 seconds with Karpenter), container image pull (30 to 60 seconds for inference images), and model loading (45 to 90 seconds), totaling 3 to 4 minutes. Teams targeting sub-200ms P50 latency cannot use scale-to-zero for user-facing traffic but can use it for batch inference, CI/CD model evaluation, and shadow deployments.
| Cost Strategy | GPU Cost/GPU/hr | Savings vs On-Demand | Risk | Best For |
|---|---|---|---|---|
| On-demand (reserved) | $3.50 (H200) | 0% | None | Production critical |
| Spot (H200, US East) | $1.50 | 57% | 8-15% interruption | Batch, dev, multi-replica |
| Spot (B200, US East) | $3.50 | 59% | 3-7% interruption | Multi-replica inference |
| Over-provision to 60% | $3.50 | 0% (extra 40% GPUs) | Higher base cost | Latency-sensitive |
| Scale-to-zero (off hours) | $0.00 (off) | 65% | Cold start delay | Dev/staging, batch |
| 1-year reserved contract | $2.80 | 20% | Capacity commitment | Stable traffic patterns |
Production Recommendations for Mid-2026
The consensus among infrastructure teams running K8s GPU inference at scale in mid-2026 converges on a reference architecture. Use Karpenter as the cluster autoscaler with at least 3 GPU instance types across 2 availability zones to handle spot interruptions without manual intervention. Run vLLM for all single-model LLM inference and Triton for multi-model or MIG-partitioned deployments. Configure HPA with queue depth as the primary scaling signal, GPU utilization as a secondary signal, and a 5-minute scale-down cooldown. Reserve at least 30% cluster headroom for spot failover capacity. Add a PodDisruptionBudget with minAvailable=1 for every inference deployment to prevent all replicas being evicted during a spot interruption or node drain. For teams exceeding $200,000/month in inference GPU cost, negotiate a 1-year reserved contract covering 60% of expected capacity and cover the remaining 40% with spot instances.
