THE THREE COMPUTE STAGES OF RAG
A production RAG pipeline has three distinct GPU-relevant stages: embedding generation, vector search, and LLM generation. Embedding is throughput-bound for ingestion and latency-sensitive for queries. Vector search is memory-bandwidth-bound, dominated by HBM reads. Generation is memory-capacity-bound, constrained by KV cache per concurrent request.
The cost distribution at 1,000 QPS is: 5-10 percent embedding generation, 2-5 percent vector search, and 85-93 percent LLM generation. The dominance of generation means RAG optimization should focus on reducing generation cost through better retrieval quality, smaller context windows, and smaller models.
EMBEDDING MODEL INFERENCE ON GPU: E5, BGE, AND COHERE
Current embedding models use decoder-only transformer architectures with mean pooling. E5-Mistral-7B-instruct (7.1B) achieves top MTEB scores but requires 14 GB VRAM and 85 ms per query. BGE-M3 (567M) achieves 63.1 MTEB at 1.1 GB VRAM and 8 ms per query, offering a 10x cost advantage for bulk indexing.
For document ingestion at scale (10M+ documents), BGE-M3 on a single H100 embeds 1.5M documents per hour versus 130K for E5-Mistral-7B. INT8 quantization of E5-Mistral reduces VRAM to 7 GB and increases throughput by 2.4x with only 0.5-1.0 MTEB point drop.
| Model | Params | MTEB Avg | VRAM | H100 Docs/s | Cost/1M Docs | Lat P50 |
|---|---|---|---|---|---|---|
| BGE-base | 137M | 59.0 | 0.6 GB | 8,500 | $0.30 | 4 ms |
| BGE-M3 | 567M | 63.1 | 1.1 GB | 3,800 | $0.65 | 8 ms |
| Cohere v3 | 1.2B | 62.8 | 2.4 GB | 1,600 | $1.55 | 15 ms |
| E5-Mistral INT8 | 7.1B | 64.1 | 7.0 GB | 480 | $5.20 | 45 ms |
| E5-Mistral FP16 | 7.1B | 65.0 | 14 GB | 290 | $8.60 | 85 ms |
| Voyage-2 | ~500M | 64.2 | 1.0 GB | 3,200 | $0.78 | 10 ms |
GPU-ACCELERATED VECTOR DATABASES: RAFT AND CUDA-OPTIMIZED INDEXING
Vector search is a memory-bandwidth problem. CPUs struggle because DDR5 bandwidth (30-50 GB/s) is 50-100x lower than HBM3 (3.35 TB/s). GPU-accelerated search using NVIDIA RAFT (cuVS) achieves 10-50x speedup over CPU FAISS. RAFT's CAGRA algorithm achieves 99.9 percentile recall at 10,000 QPS on a single H100 for 1M vectors at 768 dimensions.
The practical deployment pattern uses a hybrid index: base index in CPU DRAM with scalar quantization and a small GPU-resident hot index in full precision. This achieves P99 latency of 5-8 ms for the hot set and 15-25 ms for the full set on a 10M-vector corpus.
The L40S with 864 GB/s HBM delivers approximately 65 percent of the H100's search throughput at 30 percent of the cost, making it the most cost-effective option for GPU-accelerated vector search.
| Config | CPU FAISS | CPU FAISS+PQ | GPU RAFT L40S | GPU RAFT H100 |
|---|---|---|---|---|
| QPS (1M@768d) | 1,200 | 8,500 | 28,000 | 42,000 |
| P99 Latency | 28 ms | 6 ms | 3 ms | 2 ms |
| Recall@10 | 0.99 | 0.92 | 0.99 | 0.99 |
| Memory | 32 GB DRAM | 8 GB DRAM | 16 GB VRAM | 16 GB VRAM |
| Cost per QPS | $0.033 | $0.004 | $0.002 | $0.003 |
HYBRID SEARCH AND CROSS-ENCODER RE-RANKING ON GPU
Production RAG combines dense vector similarity with sparse keyword matching (BM25). The combined set is re-ranked by a cross-encoder model on GPU. A MiniLM-L12 cross-encoder (66M) processes 50 query-document pairs in 8-12 ms on H100.
The re-ranking stage adds $0.002-0.005 per query in GPU cost for 100 candidates. A two-stage re-ranker uses a small model for the initial pass (top 200 to top 50), then a larger model for final ranking (top 50 to top 10), reducing latency by 40 percent with identical final quality.
The key optimization is batch re-ranking across multiple concurrent queries: with 4 queries at 50 candidates each, the GPU processes 200 pairs in a single batch, achieving 4.5x throughput over serial processing.
GENERATION STAGE: SMALL LLMS AND CONTEXT MANAGEMENT
A Llama-3.1-8B generating a 200-token answer with 2,000 tokens of context costs $0.0025 per query and achieves 65-70 percent answer accuracy on multi-hop QA. Increasing context to 32,000 tokens increases cost to $0.012 per query while improving accuracy by only 3-5 percent.
At 100 QPS with 8K context: 4x Llama-3.1-8B on H100 GPUs handles the load at 75-85 percent utilization. At 1,000 QPS: 35-40 H100 GPUs. A cost-optimized architecture uses a 3B-8B model for 90 percent of queries with a 70B model for the remaining 10 percent, reducing total GPU count by 60-70 percent.
MONITORING, EVALUATION, AND GPU UTILIZATION FOR RAG
RAG monitoring spans three stages. Embedding stage: throughput (docs/s), VRAM utilization (60-80 percent target), and quantization drift. Vector search stage: P50/P99/P99.9 latency, recall@k, and GPU cache hit rate. Generation stage: tokens per second, time-to-first-token, and factual consistency score from an LLM-as-judge eval.
Inference GPUs experience bursty utilization: 30-50 percent average with spikes to 90+ percent. Effective utilization of 40-60 percent is typical, with the remainder lost to waiting and batching inefficiency. Kubernetes with GPU-informatic autoscaling can improve utilization to 60-75 percent.
Nightly retrieval evaluation with 10,000 ground-truth queries costs approximately $5-15 per day in GPU compute, a worthwhile investment to maintain RAG quality at scale.
