The KV Cache Memory Wall
The KV cache is the dominant memory consumer in LLM inference at scale. For a 70B parameter model with 8k token context and FP16 storage, a single sequence consumes roughly 2.6 GB of GPU VRAM just for the key-value cache. At batch size 64 with 4k sequences each, the KV cache alone can exceed 160 GB -- more than the HBM capacity of a single H200 GPU.
As context windows grow to 128k, 1M, or more tokens, the KV cache scales linearly with sequence length. A Llama 4 Behemoth-class model serving 1M token contexts at batch size 8 requires over 1.2 TB of KV cache storage. Without aggressive compression, serving long-context LLMs demands GPU counts that make inference economically prohibitive.
KV Cache Anatomy and Quantization Targets
The KV cache stores the key and value tensors from each attention layer across all tokens in the sequence. For a transformer with L layers, H attention heads, D head dimension, and T tokens, the cache size is 2 * L * H * D * T * bytes_per_element. The key nuance: different attention heads exhibit different sensitivity to quantization, and the value cache is generally more tolerant of compression than the key cache.
Layer-wise profiling reveals that early layers (first 20-30%) tolerate INT4 quantization with under 0.5% accuracy degradation, while middle and later layers require INT8 or FP8 to maintain output quality. This non-uniform sensitivity is the foundation for mixed-precision KV cache schemes like KIVI and KVQuant, which allocate different precisions per layer or per head.
| Precision | Bytes/Element | 70B 8k Cache Size | Accuracy (vs FP16) |
|---|---|---|---|
| FP16 | 2 | 2.6 GB/seq | Baseline |
| FP8 | 1 | 1.3 GB/seq | -0.1% |
| INT8 | 1 | 1.3 GB/seq | -0.2% |
| INT4 | 0.5 | 0.65 GB/seq | -0.8% |
| NF4 | 0.5 | 0.65 GB/seq | -0.5% |
| Mixed (KIVI) | 0.6 avg | 0.78 GB/seq | -0.3% |
FP8 KV Cache: The Blackwell Native Path
NVIDIA's Blackwell architecture (B200, B300) introduces native FP8 transformer engine support that extends to the KV cache path. With FP8 quantization, KV cache memory consumption halves versus FP16 with negligible accuracy loss for models up to 70B parameters. The B300's Transformer Engine handles the FP8 compute path in the attention mechanism directly, avoiding the dequantization overhead that plagues software-only approaches.
The effective throughput gain comes from two sources: reduced memory pressure allows larger batch sizes, and FP8 attention compute runs at roughly 1.8x the throughput of FP16 on Blackwell Tensor Cores. For batch-size-limited workloads like long-context inference, the combination delivers up to 2.5x more tokens per second versus FP16 cache on the same hardware.
INT4 and NF4: Maximum Compression with Calibration
INT4 KV cache quantization compresses the cache by 4x versus FP16, but requires careful calibration to avoid accuracy collapse. The optimal approach uses per-channel quantization with group sizes of 32-128 elements and dynamic range clipping based on activation statistics collected during a calibration run. Without calibration, naive INT4 quantization can degrade downstream accuracy by 2-4% on MMLU and HumanEval.
NF4 (Normal Float 4), introduced by QLoRA and adopted in the KV cache context by frameworks like AWQ and GPTQ, uses a non-uniform quantization grid that allocates more representational power to values near zero. The KV cache value distribution is approximately normal with a heavy tail, making NF4 a better fit than uniform INT4. In practice, NF4 achieves INT4-level compression with INT8-level accuracy -- the sweet spot for production deployments.
KIVI: Non-Uniform Per-Layer Quantization
KIVI (from the paper KIVI: A Tuning-Free Asymmetric 2-bit Quantization for KV Cache) is the most practical open-source framework for mixed-precision KV cache quantization. It assigns per-layer quantization precisions based on the layer's sensitivity to compression. Early layers use 2-bit, middle layers 4-bit, and the final layers remain at 8-bit. The key insight is that attention entropy varies significantly across layers, and low-entropy layers tolerate aggressive quantization.
In production tests on Llama 3.1 70B with 32k context, KIVI achieves a 4.2x compression ratio with less than 0.5% accuracy degradation on standard benchmarks. The implementation is tuning-free -- it uses a one-shot calibration pass on 256 sequences to determine per-layer sensitivity. KIVI is supported in vLLM and TensorRT-LLM as of mid-2026, making it deployable without custom kernel development.
| Model | FP16 Cache | KIVI Cache | Speedup | Accuracy Delta |
|---|---|---|---|---|
| Llama 3.1 8B | 4.1 GB | 0.98 GB | 1.4x | -0.2% |
| Llama 3.1 70B | 28.7 GB | 6.8 GB | 1.8x | -0.4% |
| Qwen 2.5 72B | 29.5 GB | 7.0 GB | 1.7x | -0.3% |
| Mixtral 8x22B | 22.4 GB | 5.3 GB | 1.5x | -0.5% |
Hardware Implications: VRAM Constraints and GPU TCO
KV cache compression directly changes the GPU capacity equation. For example, serving Llama 3.1 70B with 128k context and batch size 32 requires roughly 680 GB of KV cache storage at FP16. At 4x compression, this drops to 170 GB, allowing a single H200 node (8x141 GB = 1,128 GB total) to handle the workload where FP16 would require 3-4 nodes. The node count reduction from 4 to 1 translates to roughly 75% lower inference cost per token.
The choice of quantization method should be driven by the GPU generation you deploy on. On Hopper GPUs (H100, H200), INT8 and KIVI are the optimal choices since FP8 cache support is limited. On Blackwell (B200, B300), FP8 native cache is the default path with minimal integration effort, while INT4/NF4 provides additional headroom for extreme context lengths. The cost differential between INT8 and FP8 cache deployments on H200 is approximately 15-20% in favor of INT8, but the engineering effort to implement and validate INT8 pipelines often offsets the savings for smaller teams.
Implementation Guide for Production Deployments
The recommended deployment path depends on your inference framework. For vLLM users, built-in FP8 KV cache support is available via the kv_cache_dtype parameter. Enable with --kv-cache-dtype fp8 on Blackwell or --kv-cache-dtype int8 on Hopper. For KIVI integration, use the kivi-kv-cache branch of vLLM (merged upstream) with --kv-cache-dtype kivi and a calibration dataset path.
For TensorRT-LLM deployments, configure KV cache quantization in the model builder config with kv_cache_quant_mode set to FP8 or INT8. TensorRT-LLM also supports per-token dynamic quantization which adapts the quantization scale per token rather than per sequence, capturing the variance in attention patterns across a generation. This adds roughly 3% overhead to the attention computation but improves accuracy by 0.2-0.4% on long generations.
Recommendations by Deployment Profile
For high-throughput short-context serving (batch sizes above 128, context under 8k), FP8 KV cache on Blackwell GPUs provides the best accuracy-to-compression ratio with minimal engineering overhead. The memory savings enable larger batch sizes, which directly improves throughput in the compute-bound regime.
For long-context serving (32k-1M context), KIVI mixed-precision quantization on either Hopper or Blackwell GPUs delivers the best compression ratios with acceptable accuracy. The 4x compression enables workloads that would otherwise require multi-node configurations to fit on a single node, dramatically reducing latency and cost.
For cost-sensitive deployments where model accuracy is the highest priority, INT8 cache on Hopper or FP8 on Blackwell provides a safe 2x compression with sub-0.2% accuracy degradation. The implementation effort is minimal and the risk profile is well understood across production deployments.
