vLLM Overview
vLLM (vLLM 0.6) by UC Berkeley provides GPU-accelerated inference with features: PagedAttention, continuous batching, prefix caching, multi-LoRA, FP8, AWQ, GPTQ, tensor parallelism... It achieves 95-98% GPU utilization under load, 20-50ms prefill (2K tokens) throughput on H100 GPUs. License: Open-source, MIT, 50K+ GitHub stars.
Performance Benchmarks
On H100 80GB with Llama 4 Scout (17B) at FP8: prefill throughput: 21,861 tokens/second; decode throughput: 3653 tokens/second per user with 2008 max batch; TTFT (time to first token): 43ms; inter-token latency: 20ms. GPU utilization: 95-98% GPU utilization under load.
Feature Comparison
Key features: PagedAttention, continuous batching, prefix caching, multi-LoRA, FP8, AWQ, GPTQ, tensor parallelism... Unique strengths: PagedAttention continuous batching prefix caching. Production features include: OpenAI-compatible API, streaming, tool calling.
Cost-Per-Token Analysis
Cost-per-token on H100 80GB at $2.50/hr: input tokens: $0.000023/1K tokens; output tokens: $0.000089/1K tokens with vLLM. At 50% utilization, cost-per-million tokens: $34-$346 for output tokens, depending on batch size and model size. Reserved pricing reduces costs by 30-50%.
Production Deployment
Deploy vLLM in production: containerized deployment with Docker + NVIDIA Container Toolkit; Kubernetes with GPU node pools; monitoring with Prometheus + GPU metrics; horizontal scaling with Istio service mesh; and CI/CD integration for model updates. Recommended: 6x H100/B200 GPUs per node with NVLink.
When to Choose vLLM
Choose vLLM when: PagedAttention continuous batching are critical for your workloads; Open-source, MIT, 50K+ GitHub stars license model fits your budget; and your team has experience with UC Berkeley's ecosystem. Consider alternatives when: specific hardware optimization is needed, team familiarity with other frameworks, or license costs are prohibitive for your scale.
