How Prefix Caching Works
Prefix caching exploits a simple observation: in chat applications, agent loops, and RAG pipelines, large portions of the input context repeat across requests. The system prompt, tool definitions, few-shot examples, and retrieved document prefixes are identical from one request to the next. Without prefix caching, every new request reruns the full prefill computation for these repeated tokens, consuming GPU compute and HBM bandwidth on work already done.
When prefix caching is enabled, the inference engine stores the KV cache entries for common prefixes in GPU memory or host memory. On a subsequent request with a matching prefix, the engine loads the cached KV entries and only computes attention for the new suffix tokens. This shifts the compute profile from O(input_length) to O(cached_prefix) + O(new_suffix), which is dramatically cheaper for high-cache-hit workloads.
Implementation Landscape
The two major implementations are vLLM's Automatic Prefix Caching (APC) and SGLang's RadixAttention. Both use a radix tree data structure to index cached prefixes, but they differ in eviction policy and GPU memory management. vLLM's APC operates at the block level (16 tokens per block by default) and caches KV blocks in GPU memory, evicting least-recently-used blocks when memory pressure triggers.
SGLang's RadixAttention uses a finer-grained token-level radix tree that supports prefix sharing across requests even when prefixes differ mid-sequence. This yields 5-15% higher cache hit rates in workloads with variable system prompts, at the cost of slightly higher bookkeeping overhead. SGLang also supports offloading cached prefixes to host DRAM via the prefix cache store, extending effective cache capacity beyond GPU memory limits.
Measured Cost Savings by Workload
The savings vary dramatically by workload type. Chat applications with long system prompts and short user messages see the highest hit rates. Agent workloads with tool-calling loops where the conversation history and tool definitions repeat across turns also benefit significantly. RAG workloads see moderate savings because the retrieved document chunks change between queries.
In production deployments on H200 clusters, teams report prefill FLOPs reduction of 50-70% for chat, 40-55% for agent workloads, and 20-35% for RAG. The effective throughput gain is lower because decode stage costs are unchanged, but total latency per request drops proportionally to the prefill reduction. End-to-end GPU-hour savings range from 30-60% depending on average context length and cache hit ratio.
| Workload Type | Prefill FLOPs Reduction | Total GPU-Hour Savings |
|---|---|---|
| Chat (long system prompt) | 50-70% | 45-60% |
| Agent (tool loops, history) | 40-55% | 35-50% |
| RAG (variable retrieved docs) | 20-35% | 15-30% |
| Code completion (IDE) | 55-75% | 40-55% |
| Batch classification | 60-80% | 50-65% |
Memory and Throughput Tradeoffs
Prefix caching consumes GPU memory that could otherwise serve additional concurrent requests. On an H200 with 141GB HBM3e, storing KV cache for a 128K-token prefix of a Llama 4 Maverick model consumes approximately 24GB. Serving 1,000 unique prefixes simultaneously would require more HBM than available, forcing a tradeoff between cache diversity and batch size.
The optimal configuration depends on prefix diversity. Workloads with 10-50 common prefixes benefit from GPU-side caching with LRU eviction. Workloads with hundreds or thousands of prefixes should use host-side caching with GPU prefetch on cache hit, accepting 2-5ms of PCIe transfer latency per hit in exchange for dramatically larger cache capacity. SGLang's prefix cache store supports this tiered approach natively.
TCO Impact at Cluster Scale
At cluster scale, prefix caching directly reduces the number of GPUs needed to serve a given request rate. For a production chat application serving 10K requests per minute with average 4K-token prefixes and 500-token generations, a cluster without prefix caching requires approximately 32 H200 GPUs to meet latency SLOs. With prefix caching achieving a 60% hit rate, the same throughput requires 18-20 H200 GPUs.
The annual cost difference is roughly $600,000-800,000 at current spot rates of $3.10/GPU/hr. This is pure infrastructure savings with no model quality degradation, making prefix caching the single highest-ROI optimization available for production LLM serving today.
Implementation Best Practices
Enable Automatic Prefix Caching in vLLM by setting `--enable-prefix-caching` or configure RadixAttention in SGLang with appropriate cache size limits. Start with GPU-only caching set to 30-50% of available HBM, then introduce host-side tiered caching once hit rates stabilize. Monitor the `avg_prefix_cache_hit_rate` metric closely during rollout.
Optimize system prompt design for cacheability. Place static content (tool definitions, instructions, safety guardrails) at the start of the prompt and variable content (user messages, retrieved documents) at the end. Tools like LangChain and DSPy can automatically segregate prompts into cacheable and non-cacheable segments. Avoid putting user-specific data in the first 1,000 tokens.
Our Recommendation
Every production LLM serving deployment should implement prefix caching. It is free in terms of model quality, requires no model modification, and delivers 30-60% GPU cost reduction in the workloads that dominate current production traffic. The implementation effort is measured in days using vLLM or SGLang, not weeks.
For teams building new inference stacks, start with SGLang's RadixAttention for its superior prefix sharing and tiered cache support. For existing vLLM deployments, enable APC and tune the max prefix length and block size. Monitor cache hit rates and adjust memory allocation. The ROI is immediate and measurable in GPU-hour consumption.
