All essays
TechnicalDEEP DIVEFEB 2026

GPU Inference Caching Strategies: Semantic Cache, KV Cache, Response Cache

Deep dive into LLM inference caching tiers: prefix KV cache reuse, semantic embedding cache, and response cache. Latency reduction, hit rate optimization, and GPU cost savings for production serving on H100 and A100 clusters.

01

THE THREE-TIER INFERENCE CACHE HIERARCHY

LLM inference is memory-bandwidth bound, not compute bound, for most production serving configurations. A Qwen 2.5 72B model with 4K sequence length on an H100 SXM spends roughly 75% of inference time on attention operations dominated by KV cache reads. Caching strategies address this at two levels: reusing KV computations across requests with shared prefixes, and avoiding redundant LLM calls entirely through semantic or exact-match result caching. A serving stack processing 10 million tokens per hour with 35% effective cache reduction reclaims 3.5M tokens of GPU time per hour, saving roughly $3,200 per month on an 8-GPU H100 deployment at $1.15/hr per GPU.

The three caching tiers are: Tier 1 - prefix KV cache reuse within the LLM engine (vLLM, TensorRT-LLM) operating at GPU memory level with zero additional latency; Tier 2 - exact-match application cache (Redis, Memcached) keyed on input token hash, catching identical requests processed by different engine instances; Tier 3 - semantic embedding cache using vector similarity to detect semantically equivalent queries, catching the largest fraction of redundant traffic at the cost of 5-15ms additional latency for embedding computation. Each tier addresses a different failure mode in the cache hierarchy with distinct cost and latency tradeoffs.

Cache TierMatch CriterionTypical Hit Rate
Tier 1: Prefix KV CacheExact token prefix20-40%
Tier 2: Exact-Match CacheSHA256 of full input15-30%
Tier 3: Semantic CacheCosine similarity > 0.9530-60%
Combined Three-TierProgressive fallthrough50-75%
Embedding Latency Cost5-15ms per lookupNegligible vs LLM gen
02

PREFIX KV CACHE: INSIDE THE LLM ENGINE

vLLM implements prefix-aware KV cache management through automatic prefix caching. When enabled, PagedAttention stores KV blocks in a hash table indexed by the hash of their token sequence. Requests arriving with tokens matching a cached prefix automatically reuse the cached KV blocks, only computing new KV pairs for the suffix. The cache eviction follows an LRU policy within each GPU's memory budget. The hash lookup completes in under 1 microsecond and requires no inter-GPU coordination, making this essentially free performance when prefix overlap exists in the traffic pattern.

The memory impact is significant. For a 70B model at FP16 with 4K sequence length on an 8x H100 node, the KV cache consumes approximately 30 GB per GPU without prefix reuse. A 500-token shared system prompt that is fully cached saves roughly 4 GB per GPU of KV cache memory. With serving batch sizes of 16-32 concurrent requests, the memory savings compound directly into higher throughput per GPU. Prefix caching is most effective for chatbot deployments with fixed system prompts, RAG pipelines with shared instruction prefixes, and few-shot scenarios where the same examples preface every query. vLLM's prefix cache also supports partial matches: if a request shares only 200 of 500 prefix tokens, those 200 tokens are still reused.

TensorRT-LLM implements prefix caching through its KV reuse API, which allows explicit management of cached prefix blocks. Unlike vLLM's automatic approach, TensorRT-LLM requires the application layer to identify cached prefixes and pass cache hint metadata to the inference runtime. This gives more control at the cost of integration complexity. TGI does not natively support KV cache reuse - its contiguous KV cache design makes block-level prefix sharing difficult without full re-implementation of the cache manager.

03

SEMANTIC CACHE: EMBEDDING-SIMILARITY QUERY MATCHING

Semantic caching extends beyond exact token matching by using embedding similarity to detect when different queries express the same intent. Two queries like "What is your refund policy?" and "Can I get my money back?" have different token sequences but the same answer. A semantic cache embeds each incoming query into a vector, performs an approximate nearest neighbor search against stored query-response pairs, and returns the cached response if similarity exceeds a configurable threshold. The embedding computation adds 5-15ms, compared to 500-5,000ms saved by avoiding the LLM call.

Implementation requires an embedding model - Sentence-BERT, E5-mistral-7b, or the embedding endpoint of your LLM provider - and a vector store (Milvus, Qdrant, PGVector, or FAISS index). The threshold selection is the most critical tuning parameter. At 0.97 cosine similarity, the cache is conservative with near-zero false positives but hit rates of 20-35%. At 0.92, hit rates rise to 40-55% but false positive risk increases. For customer support applications where incorrect cached responses have high reputational cost, start at 0.97 and tune downward with spot-check audits. For documentation Q&A with well-defined answers, 0.92-0.95 is appropriate. Track cache correctness via a background job that periodically reruns a sample of cached queries through the LLM and compares responses.

04

RESPONSE CACHE: EXACT-MATCH WITH TTL STRATEGIES

The response cache is the simplest tier: a key-value store (Redis, memcached) mapping an input hash to the generated response. It catches identical requests that arrive at different server instances or at different times. In production, 15-30% of inference traffic is typically exact duplicates - retries from client-side backoff, polling patterns, identical queries from multiple users within a short window, or automated systems sending the same request periodically. The response cache returns cached results in under 1ms, essentially zero incremental latency.

TTL strategy determines cache freshness. For dynamic knowledge applications (news, product information, pricing), a TTL of 5-15 minutes balances freshness with cache efficiency. For static knowledge (documentation, policies, FAQ), a TTL of 24-72 hours is appropriate. The most sophisticated approach is active invalidation: when underlying data changes, push invalidation events to the cache keyed on query patterns. For example, if a product page updates, invalidate all cache entries whose embedding vectors cluster near queries containing that product name. Combined with semantic cache, this hybrid approach delivers cumulative hit rates of 50-75% depending on query diversity.

05

CROSS-INSTANCE DISTRIBUTED CACHING

When scaling beyond a single GPU node, prefix caching becomes fragmented. Each GPU manages its prefix cache independently, so a request hitting GPU B cannot reuse a prefix computed on GPU A. For clusters serving a common set of prompts across multiple replicas, this fragmentation wastes significant compute. The solution is a distributed KV cache that pools prefix state across all GPUs using a shared backing store like Redis or Memcached, indexed by token prefix hash.

The distributed cache adds serialization-deserialization overhead of roughly 1-3ms for a 500-token prefix, compared to the 50-150ms saved on redundant computation. However, this requires fast networking between GPU nodes and the cache store. A Redis cluster on the same InfiniBand fabric reduces serialization latency to 200-400 microseconds. Cloud-hosted Redis adds 5-20ms network latency that can negate the caching benefit for short prefixes. Co-locate cache instances on the same cluster network as your GPUs, ideally within the same rack, to keep networking overhead below 1ms. Below 200-token prefixes, the distributed cache overhead exceeds the computation saved, making instance-local caching more efficient.

06

CACHE ROI: AUXILIARY COST VERSUS INFERENCE SAVINGS

Caching infrastructure adds its own costs: embedding model inference for semantic caching (a small model like all-MiniLM-L6-v2 on CPU or E5-mistral-7b on a GPU fraction), the vector database or Redis cluster, and engineering time. A full three-tier cache deployment costs roughly $300-800 per month in auxiliary compute and storage, plus initial engineering of 2-4 weeks. The payback is typically 2-6 weeks depending on inference volume.

For a cluster serving 5M tokens per day from a 70B model at $1.15/hr per GPU, daily inference cost is roughly $350 at approximately $0.37 per million tokens. A 50% cache hit rate reduces effective token volume to 2.5M tokens per day, saving $175 per day - roughly $5,250 per month. The caching infrastructure cost of $500 per month is recovered in under 3 days. The ROI flips negative only for very low-volume deployments under 100k tokens per day. Beyond GPU savings, cached responses return in 100-300ms compared to 2-8 seconds for full generation, improving user-facing metrics where every 100ms of latency reduction increases conversion by 1-2%.

Deployment ScaleMonthly Cache CostMonthly GPU Savings
Small (1M tok/day)$150-300$200-500
Medium (5M tok/day)$300-500$3,000-6,000
Large (50M tok/day)$500-1,500$30,000-60,000
X-Large (500M tok/day)$2,000-5,000$300,000-600,000
Payback Period2-6 weeksNegative after day 14-42
Filed under
KV Cache OptimizationSemantic CachingResponse CacheLLM InferenceGPU MemoryPrefix CachingInference Cost Reduction