Three GPU Workloads, Three Very Different Hardware Profiles
RAG GPU infrastructure in 2026 is not one compute problem - it is three separate workloads stitched together, each with fundamentally different hardware requirements. The mistake most teams make at prototype stage is running everything on whatever H100 cluster is available. At 1M daily queries, that mistake shows up on a monthly cloud bill. At 10M queries, it costs a full-time engineer in wasted procurement.
Embedding generation is high-throughput and memory-bandwidth-bound, but compute-light. A 568M-parameter encoder like BGE-M3 uses roughly 2-3% of an H100's tensor core capacity at typical batch sizes. The real bottleneck is tokenization, batch assembly, and the PCIe transfer overhead between CPU and GPU. You are paying for H100 flops you cannot consume. Vector search is almost entirely memory-access-bound - whether you are doing exact inner-product search in FAISS or approximate nearest-neighbor on an HNSW graph index, the GPU is doing irregular sparse memory access across your index, not dense matrix math. LLM generation is the stage that actually needs a serious GPU. Autoregressive decoding is memory-bandwidth-bound during the decode phase, compute-bound during prefill, and most production RAG queries present long prompts from retrieved context that stress both phases simultaneously.
Treating these three workloads as one GPU procurement decision is expensive. The right architecture matches each stage to its hardware profile, and the right procurement approach sources all three tiers from a single desk so you are not managing three separate provider relationships. That is the case ClusterBid was built for.
Embedding Model GPU Sizing - Why L4 Beats H100 on Cost Per Embedding
Here is something that rarely makes it into documentation: embedding models are so compute-efficient relative to their output throughput that an H100 is overkill by a factor of five. BGE-M3 at INT8 quantization processes approximately 85,000 tokens per second on a single NVIDIA L4 24GB at batch size 512. E5-large-v2 runs faster. The L4's 24GB VRAM holds BGE-M3 with comfortable headroom, and at roughly $0.42/hr on spot, the cost-per-million-embeddings math favors the L4 over H100 at every realistic query volume below 500M embeddings per day.
The practical threshold: if you need more than 300M embeddings per day (approximately 30M queries with ten chunks each), scale L4s horizontally before considering H100s. A four-L4 node costs under $1.70/hr and handles the embedding load for 15-20M daily queries with room to spare. ClusterBid sees consistent spot availability on L4 and L40S that is currently tighter on H100 SXM - at enterprise query volumes you can often source embedding GPU capacity in days rather than weeks.
One practical note on model choice: BGE-M3 is a strong default in 2026 because it handles multilingual content and long-context documents (up to 8192 tokens) in a single encoder. If you are embedding short passages in English only, E5-large-v2 or even BGE-small runs faster and cheaper. Match the model to your document length distribution before sizing hardware.
| GPU | BGE-M3 tokens/sec (INT8, bs=512) | Cost per 1M embeddings |
|---|---|---|
| L4 24GB | ~85,000 | $0.0017 |
| A10G 24GB | ~110,000 | $0.0021 |
| L40S 48GB | ~185,000 | $0.0021 |
| H100 PCIe 80GB | ~320,000 | $0.0020 |
| H100 SXM 80GB | ~390,000 | $0.0019 |
GPU-Accelerated Vector Search - The 50M Vector, 5K QPS Breakeven
For most teams building production RAG in 2026, GPU-accelerated vector search does not pay off. The threshold is roughly 50M vectors in your index with sustained query throughput above 5,000 QPS. Below that, a well-tuned CPU-based deployment of Qdrant, Milvus, or Weaviate with HNSW indexing on modern AMD EPYC or Intel Xeon hardware delivers recall and latency that GPU acceleration cannot meaningfully improve, at a fraction of the cost.
Above those thresholds, the gap opens fast. Milvus benchmarks show GPU IVF-Flat achieving around 15,000 QPS at 100M vectors compared to approximately 2,000 QPS on a 32-core CPU node - a genuine 7x advantage. RAPIDS cuVS extends this further with CAGRA, a graph-based approximate nearest-neighbor algorithm that runs 20-50x faster than HNSW on GPU at equivalent recall rates. If you are running Milvus or Qdrant with GPU indexes enabled, cuVS is the underlying acceleration layer. For organizations indexing tens of billions of document chunks - think enterprise knowledge bases or large-scale document corpora - the GPU vector search math changes entirely and dedicated GPU nodes become mandatory.
One hardware note that matters: GPU-accelerated vector search benefits from PCIe form factor GPUs, not SXM. Your vector index lives in CPU RAM and gets transferred to GPU for search via PCIe. The fast NVLink interconnect on SXM cards does nothing for this workload - you are paying a 40-50% premium for bandwidth you cannot use. An A100 PCIe or L40S PCIe is the right GPU for dedicated vector search nodes.
Co-location vs Separate Pools - Three Patterns and When Each Wins
In practice, production RAG systems converge on three deployment patterns. Pattern one: single cluster, everything co-located. This is right for prototypes and small-scale deployments under 500K queries per day. You avoid multi-cluster ops complexity and H100 MIG partitioning lets you run embedding generation in smaller slices. The downside is that you are sizing every GPU for peak LLM demand while embedding processes consume 3-5% of that capacity.
Pattern two: fully separated clusters, one per stage. Highest utilization efficiency - each cluster sized exactly to its workload. Also the highest operational complexity: separate scheduling, separate cost attribution, separate vendor relationships, separate monitoring. This pattern makes sense above 10M daily queries when the GPU cost savings justify a dedicated platform engineering investment.
Pattern three: hybrid tiered cluster with L4 or L40S nodes for embedding and reranking, and H100 or H200 nodes for LLM generation. This is what most teams land on between 1M and 10M daily queries. The key insight is that these stages are sequential per query but fully parallelized across users - embedding and generation never contend for the same GPU. A RAG query spends 30-80ms on embedding and vector search, then 2-10 seconds waiting for LLM generation. You can run completely separate GPU pools with only your request queue as the shared resource.
| Pattern | Best Scale | Key Trade-off |
|---|---|---|
| Single cluster (co-located) | < 500K queries/day | Simple ops, low GPU utilization |
| Hybrid tiered (L4 + H100) | 1M - 10M queries/day | Good efficiency, moderate complexity |
| Fully separated clusters | > 10M queries/day | Highest efficiency, highest ops burden |
Re-ranking: The 20x Compute Multiplier Most Teams Miss
Cross-encoder rerankers - BGE-reranker-v2-m3, Cohere Rerank 4, or any model that jointly encodes query and document - are 20-50x slower per token than bi-encoder embedding models. They read both the query and each candidate passage simultaneously, computing attention across the full concatenated input. A reranking pass over k=20 candidates costs 20x the compute of embedding the original query. Most teams add reranking at prototype stage without thinking through what that multiplier means at production query volumes.
The numbers: BGE-reranker-large processes approximately 9,000 document pairs per second on a single L4 at 512-token context. At 1M daily queries with k=20 reranking, you are running 20M pairs per day - 232 pairs per second on average, well inside a single L4. But at 10M daily queries, that same calculation requires 3-4 dedicated L4 GPUs purely for reranking. At 100M daily queries, you need 30+ L4 GPUs for the reranking stage alone.
The optimization lever teams overlook is k - the number of candidates passed to the reranker. Most knowledge base use cases do not benefit from k above 10-15. Every increment of k you add multiplies reranking compute proportionally. Use k=20 or k=30 when retrieval recall is critical (medical, legal, financial). For general enterprise Q&A over curated content, k=10 gives you 90% of the quality improvement at half the compute cost.
Full RAG Stack Cost at 1M, 10M, and 100M Daily Queries
Building a total GPU cost model for RAG requires treating each stage separately. These estimates assume BGE-M3 for embedding, BGE-reranker-large with k=20 for reranking, a 70B parameter LLM for generation (Llama 4 Maverick or similar), and 2026 spot GPU pricing. Vector search is excluded because it runs on CPU below 50M vectors.
At 1M daily queries: embedding requires one L4 at approximately $300/month, reranking shares that L4 with headroom to spare, LLM generation on two to four H100 PCIe nodes runs $1,500-$3,000/month depending on average output length. Total GPU spend: roughly $1,800-$3,300/month. The LLM generation dominates at 80-90% of total cost - that ratio holds at every scale.
At 10M daily queries, embedding and reranking scale to three to five L40S GPUs at around $2,200-$3,700/month. LLM generation requires 16-24 H100 SXM GPUs (two to three 8-GPU nodes) and runs $12,000-$18,000/month. At 100M daily queries, the numbers shift to 25-30 L4 GPUs for embedding ($4,000-$5,000/month) and 160-200 H100 SXM GPUs ($120,000-$150,000/month) for generation. Switching to a 13B distilled model at this query volume cuts generation cost by 70% with acceptable quality degradation for most factual retrieval use cases.
| Daily Query Scale | GPU Config (embedding + LLM) | Est. Monthly GPU Cost |
|---|---|---|
| 1M queries/day | 1x L4 + 2-4x H100 PCIe | $1,800 - $3,300 |
| 10M queries/day | 3-5x L40S + 16-24x H100 SXM | $14,200 - $21,700 |
| 100M queries/day | 25-30x L4 + 160-200x H100 SXM | $124,000 - $155,000 |
Architecture Recommendations by Query Scale
Below 1M daily queries, a single 4x H100 PCIe node handles everything. Run embedding on the same GPUs via MIG partitioning or time-slicing - it is less efficient but the operational simplicity of one cluster, one provider, one monitoring setup is worth it. At this scale, you are not optimizing infrastructure, you are learning your traffic patterns.
At 1M to 10M queries per day, the hybrid tiered architecture pays off. Two to four L40S GPUs handle embedding and reranking - the L40S is better than L4 here because its 48GB VRAM handles larger context windows in the reranker and supports Llama models up to 13B for lightweight query classification tasks without a separate node. A dedicated H100 SXM cluster handles generation. ClusterBid can quote both tiers as a package - L40S for embedding, H100 SXM for generation - from the same sourcing desk, which matters when you need to negotiate interconnect requirements and rack placement for sub-100ms end-to-end latency.
Above 10M daily queries, model selection matters more than GPU selection. A 13B distilled LLM serving production RAG queries costs 70-80% less than a 70B model in GPU compute, and for knowledge retrieval tasks on curated content, the quality gap is smaller than most teams expect. Size your generation cluster to the model you actually need, not the largest model that fits. The embedding and vector search tiers are not your cost problem at this scale - they are rounding errors on your LLM generation bill.
