All essays
TechnicalDEEP DIVEFEB 2026

The Production Open-Source Inference Stack in 2026: vLLM, Ray Serve, and Prometheus for Under $15K/Month

Deploy production LLM serving open source in 2026 with vLLM Ray Serve production deployment on 4x H100 for under $15K/month. Full stack recipe with real cost breakdown.

01

The Stack: Why These 4 Components and Nothing Else

Self-hosted LLM inference under $15K/month is not about finding the cheapest components. It's about picking the narrowest possible stack that covers production requirements and then refusing to add anything else. The stack that actually works in 2026: vLLM V1 as the inference engine, Ray Serve for deployment and autoscaling, Prometheus plus Grafana for observability, and NGINX at the edge for rate limiting and TLS termination. That's it.

vLLM handles the hard inference problems - PagedAttention for KV cache management, continuous batching to keep GPU utilization above 70%, and speculative decoding for latency reduction. The V1 engine that shipped in late 2024 and matured through 2025 is meaningfully better than earlier releases: the new scheduler separates request prefill and decode phases cleanly, which matters enormously once you're running mixed workloads at production load. Ray Serve adds the HTTP serving layer, replica management, and autoscaling logic without requiring you to write any of that yourself. The production LLM serving open source story in 2026 starts here.

The thing nobody tells you about this combination is that the Ray head node is not where your model lives - it's a coordinator process that can run on a $50/month VM. Your GPU workers register with it and Ray Serve routes requests to whichever replica has capacity. This means you can add or remove GPU nodes without touching your serving config. NGINX sits at the front and handles rate limiting per API key, TLS termination, and basic request logging before anything hits your Ray cluster. Grafana reads from Prometheus. The only hard dependency you have is a way to get model weights onto your GPU nodes when they boot, which S3-compatible object storage handles for a few hundred dollars a month.

02

What $15K/Month Actually Buys in GPU Compute Capacity

For production LLM serving with a $15K/month total budget, you're choosing between two sensible configurations: 4x H100 SXM5 80GB nodes or 6x L40S 48GB nodes. Both are available on the spot and reserved markets right now. H100 SXM5 reserved rates have dropped significantly through 2026 as the enterprise upgrade supercycle pushed older Hopper inventory into the market - you can source 4x H100 SXM5 at roughly $2.00-2.50/GPU/hr on 3-month reserved contracts, which works out to $5,760-7,200/month just for GPU compute.

The L40S route is more nuanced. At 48GB VRAM per card, L40S handles Llama 4 Scout (109B MoE, ~37B active) in FP8 at 4-way tensor parallelism, or Qwen 3 72B in BF16 at 2-way tensor parallelism. L40S spot rates run $0.80-1.20/GPU/hr, and 6 cards give you $3,456-5,184/month in GPU cost with more headroom in your total budget for egress and ops tooling. The tradeoff: L40S has no NVLink, so tensor parallelism runs over PCIe - expect 15-20% throughput reduction versus H100 SXM5 for multi-GPU tensor parallel inference. For single-GPU or 2-GPU deployments serving models under 40B parameters, L40S is genuinely the better economics call.

For serving Llama 4 Maverick (400B MoE, ~52B active parameters) or DeepSeek V4, you need 8x H100 SXM5 minimum at FP8, which blows the $15K monthly budget on GPU costs alone. That's not a budget the $15K stack is designed to serve. If you're running anything in the 400B+ class, this article isn't your operating manual - you need at least 8x H200 or 4x B200, and the monthly cost starts at $18-25K just for the hardware. The $15K/month target is the sweet spot for 7B-70B dense models and medium-sized MoE models with active parameter counts under 60B.

ConfigurationGPU Cost/MonthMax Model Size
4x H100 SXM5 (reserved)$5,760-7,20070B FP8, 34B BF16
6x L40S (spot)$3,456-5,18472B FP8 (2-way TP)
8x H100 SXM5 (reserved)$11,520-14,400400B MoE FP8
4x L40S + 2x H100 (mixed)$4,800-6,60034B BF16 dual-stack
03

Deployment Architecture That Doesn't Break at 3 AM

The Ray Serve deployment config is where most teams make their first expensive mistake. The default settings assume you want to maximize throughput at the cost of latency - that's correct for batch workloads but wrong for interactive inference. For production serving under the $15K/month self-hosted LLM inference stack, set max_ongoing_requests to roughly 40-60 per replica (lower for longer context windows), configure autoscaling_config with target_num_ongoing_requests_per_replica at 20-30, and set downscale_delay_s to 300 seconds minimum so you're not constantly spinning replicas up and down during normal traffic fluctuation. The upscale delay should be near zero - you'd rather pay for 5 minutes of an extra replica than queue requests during a traffic spike.

Model loading is the silent killer of your cold-start time. Loading a 70B parameter model from S3 to GPU VRAM takes 8-15 minutes on a fresh node, depending on your network throughput. Two approaches work: either keep at least one warm replica running at all times (pay for idle GPU time but get instant scaling), or pre-load model weights to local NVMe on your GPU nodes and pull from disk instead of object storage (15-30 second load times vs 8-15 minutes). For the $15K/month budget, one always-on replica serving traffic with autoscaling for bursts is the right call - you're paying $48-60/day for that baseline GPU but eliminating the cold-start problem entirely.

Tensor parallelism configuration in vLLM deserves more attention than it typically gets. For H100 SXM5, set tensor_parallel_size=2 for 34B models and tensor_parallel_size=4 for 70B models at BF16. At FP8, a single 80GB H100 can hold a 70B model (70B * 2 bytes FP16 = ~140GB weight-only, but FP8 halves that to ~70GB, plus KV cache overhead makes this tight - in practice, FP8 70B on a single H100 80GB works with a modest max_model_len). Use --max-model-len 32768 or lower unless you specifically need 128K context, because longer context windows eat KV cache proportionally and reduce your effective batch size. The difference between max_model_len 8192 and 128K on batch throughput is roughly 10x.

04

5 Dashboards You Actually Need (Not 15 You'll Never Look At)

The production LLM serving open source observability story in 2026 starts with DCGM (Data Center GPU Manager). NVIDIA's dcgm-exporter runs as a DaemonSet on each GPU node and exports hardware metrics to Prometheus: GPU utilization per card, memory utilization, power draw, NVLink bandwidth, and temperature. These four metrics - utilization, memory, power, and NVLink - tell you immediately whether a GPU is underloaded, memory-bound, power-capped, or experiencing interconnect bottlenecks. Add vLLM's built-in Prometheus metrics endpoint (it exposes request latency, queue depth, KV cache utilization, and token throughput by default) and you have the full picture.

Dashboard 1: GPU Utilization Panel. Target: all serving GPUs above 60% utilization during traffic hours. Below 40% means you're over-provisioned or your autoscaling is too aggressive. Dashboard 2: Request queue depth and TTFT (time to first token) p50/p95/p99. TTFT p95 above 3 seconds for interactive use cases is your PagerDuty trigger. Dashboard 3: Token throughput (tokens/second) and cost per million tokens - divide your hourly GPU cost by throughput to get real-time cost efficiency. Dashboard 4: KV cache utilization. vLLM exposes vllm:gpu_cache_usage_perc - when this consistently runs above 85%, you're getting close to out-of-memory errors or forced request evictions. Dashboard 5: HTTP error rate and 5xx distribution by model endpoint. If one model version is throwing errors and another isn't, you catch it in seconds instead of waiting for a user complaint.

Alert rules that have saved production systems: (1) PagerDuty on TTFT p95 > 5s for 3 consecutive minutes. (2) Warning (Slack only) when KV cache utilization > 80% for 10 minutes. (3) Critical on any GPU temperature > 85C. (4) Warning when any GPU utilization drops below 20% for 30 minutes during business hours - this catches stuck processes and misconfigured autoscaling. (5) Critical when vLLM reports num_requests_waiting > 100 for more than 5 minutes - your serving capacity is saturated and users are queuing. These five rules replace a 20-rule alert matrix that most teams build and then tune-out because of false positives.

05

The Real Monthly Spend: Every Line Item Against OpenAI API Equivalent

Here's the actual $15K/month budget breakdown for a 4x H100 SXM5 vLLM Ray Serve production deployment serving Llama 4 Scout at FP8. GPU rental at reserved rates: $6,480/month (4 GPUs, $2.25/hr/GPU, 720 hours). Object storage for model weights and checkpoints (200GB for a quantized 70B model): ~$25/month on S3-compatible storage. Egress for inference responses (assuming 50M tokens/month output at ~4 bytes/token average): ~$180/month. Ray head node coordination VM (4 vCPU, 16GB RAM): ~$80/month. NGINX load balancer instance: ~$40/month. Prometheus and Grafana on a small VM: ~$30/month. Ops time: ~4 hours/week from a senior ML infrastructure engineer at $200/hr amortized = ~$3,200/month. Total: approximately $10,035/month in direct costs, or $13,235/month including ops time.

The OpenAI API equivalent for 50M tokens/month (assuming a 2:1 input/output ratio, ~33M input and 17M output) at GPT-4o pricing ($2.50/M input, $10/M output): $82.50 + $170 = $252.50/month. Wait - that's massively cheaper than self-hosting. This is the calculation most articles get wrong. At 50M tokens/month, GPT-4o wins. The break-even for self-hosted LLM inference under $15K/month total cost versus GPT-4o is roughly 1.5-2 billion tokens per month. You need serious scale to justify the infrastructure overhead.

The honest case for self-hosting at under 1.5B tokens/month is not pure economics - it's data privacy (your prompts don't leave your network), latency control (no OpenAI rate limits or capacity constraints), model customization (you can fine-tune, you can run models not available via API), and vendor independence. If any of these matter to your use case, the $10-14K/month stack makes sense. If they don't, and you're under 1B tokens/month, stay on the API. Teams that move to self-hosted LLM inference because they think it's cheaper at moderate scale are almost always wrong. Teams that move because they need control are usually right.

Cost CategoryMonthly CostNotes
4x H100 SXM5 GPU rental (reserved)$6,480$2.25/GPU/hr, 3-month contract
Object storage (model weights)$25200GB S3-compatible
Egress (50M output tokens)$180~4 bytes/token avg
Ray head node VM$804 vCPU, 16GB RAM
NGINX + Prometheus VMs$70Two small instances
Ops time (4 hrs/wk senior eng)$3,200~$200/hr blended rate
Total (with ops)$10,035-13,235Direct + labor
OpenAI GPT-4o equivalent$252-2,52050M-500M tokens
06

Scaling Beyond 4 GPUs Without Rewriting Everything

The vLLM Ray Serve production deployment architecture described here scales horizontally without any architectural changes up to about 16x H100 on a single Ray cluster. Adding a 5th GPU node means registering it with the existing Ray cluster (one command), updating your autoscaling config to allow more max replicas, and optionally running a second vLLM model instance for redundancy. The Ray head node handles all scheduling. You don't change your NGINX config, your Prometheus config, or your application code.

Beyond 16 GPUs on a single cluster, you're looking at running multiple Ray clusters with a shared NGINX routing layer. This is where you'd segment by model (cluster A serves Qwen 3 72B, cluster B serves Llama 4 Scout), or by traffic tier (high-priority traffic to reserved instances, batch/async traffic to spot instances with checkpoint-based fault tolerance). The observability stack scales naturally - just add more scrape targets to your Prometheus config and update your Grafana dashboards with cluster-dimension filters.

One thing to get right early: model versioning and deployment strategy. Blue-green deployments for LLM inference work surprisingly well with Ray Serve - you start a new deployment with a new model version, shift 10% of traffic to it using Ray's traffic splitting configuration, watch your Grafana quality metrics for 30-60 minutes, then shift 100% and kill the old deployment. The GPU RAM holds both versions temporarily (this requires excess capacity), but for a 5-10 minute cutover window, the cost is trivial. ClusterBid's inventory has H100 SXM5 nodes available for on-demand spot capacity to cover blue-green transition periods - it's the kind of short-duration use case spot markets handle well.

07

Sourcing the Hardware: What to Verify Before You Sign a Contract

For the 4x H100 SXM5 configuration that anchors this stack, the two things that matter most are NVLink connectivity and NUMA topology. SXM5 H100s on a proper HGX board have NVLink 4 connecting all 8 GPUs on the server - if you're renting a 4-GPU allocation, verify you're getting 4 GPUs from the same NVLink domain, not 4 random GPUs from different nodes. Getting GPUs from different NVLink domains for a tensor-parallel inference deployment will throttle your cross-GPU bandwidth to PCIe speeds (~64 GB/s vs 900 GB/s NVLink) and tank your throughput.

SLA verification for inference workloads is different from training workloads. You need network uptime guarantees (inference is latency-sensitive, a 30-second network partition causes visible user impact) and you need a commitment on GPU replacement time if a card fails (training can checkpoint and restart; inference serving cannot - a failed GPU in a tensor-parallel group takes the replica offline). Get a 15-minute hardware replacement SLA in writing, or architect for it explicitly by running replicas across separate physical nodes so a single node failure drops capacity but doesn't take down the service.

The secondary market for H100 SXM5 servers has matured significantly through 2026 as hyperscaler upgrade cycles push Hopper hardware to secondary buyers. If your team has the operational capacity to manage physical hardware, buying used H100 SXM5 servers in the $25-35K range per 8-GPU node (prices as of mid-2026 secondary market) and colocating in a data center can bring your $15K/month operating cost down substantially versus renting. The 3-year TCO math frequently favors ownership for teams running at 80%+ utilization. ClusterBid's sourcing desk can help structure both rental and secondary market procurement - the same analysis that applies to this stack.

Filed under
vLLM V1Ray ServePrometheusSelf-Hosted LLMOpen Source InferenceH100 DeploymentDCGM Metrics