Why GPU Acceleration for Vector Search
Vector search is the backbone of retrieval-augmented generation (RAG), and it is increasingly the bottleneck in end-to-end inference pipelines. A single user query may trigger embedding of the input, ANN search over a 10M+ vector index, and reranking of the top-K results. On CPU-based systems, the ANN search step dominates latency, often taking 50–200ms per query at high recall targets.
GPU acceleration targets the distance computation at the heart of ANN search. Instead of computing cosine or L2 distances on CPU SIMD units, GPU vector databases batch distance calculations across the GPU's tensor cores, achieving 10–50x throughput improvement on the distance computation step. The wall-clock improvement is smaller (3–8x) because of PCIe transfer, index traversal, and result gathering overhead, but the cost-per-query improvement is significant at scale.
Milvus on GPU: MMap and cuVS Integration
Milvus 4.0 introduced native GPU index support through its integration with NVIDIA cuVS (formerly RAFT). The GPU IVF_FLAT and GPU IVF_PQ indexes run entirely on-GPU during search, using a single H100 to process 100K+ vectors per second at 95%+ recall. Milvus also supports a hybrid mode where the coarse quantizer runs on GPU while the residual vector search runs on CPU, balancing throughput and memory footprint.
The practical benefit: a 4-node Milvus cluster (4x H100, 64-core CPU, 256 GB RAM) serving a 50M vector index achieves 4,200 QPS at p99 latency of 35ms. The equivalent CPU-only cluster (no GPU, 128-core CPU, 512 GB RAM) achieves 1,100 QPS at p99 latency of 95ms. GPU acceleration delivers roughly 3.8x throughput at 2x node cost - a net 1.9x improvement in queries per dollar.
Qdrant with GPU Support
Qdrant introduced GPU-accelerated HNSW search in its 2.0 release using CUDA. Unlike Milvus's approach of dedicated GPU indexes, Qdrant uses GPU acceleration as an optional hardware-accelerated distance computation layer. The HNSW graph traversal still happens on CPU, but all distance calculations are offloaded to the GPU. This design reduces the memory overhead of GPU indexes (no need to store vectors in GPU memory) at the cost of higher PCIe traffic.
In benchmark testing on H100 with a 10M vector dataset (1536 dimensions), Qdrant GPU mode achieves 2.8x throughput improvement over CPU-only at the same recall target (99%). The improvement is smaller than Milvus's because HNSW graph traversal remains on CPU, but the memory efficiency means a single H100 can serve indexes of unlimited size - only the graph structure must fit in CPU RAM.
Weaviate and the cuVS Pathway
Weaviate takes a modular approach. Its vector index layer supports pluggable backends, and the NVIDIA cuVS integration is available as an experimental module. The cuVS backend offloads both index building and search to GPU, using the CAGRA (CUDA-Accelerated Graph for ANN) algorithm. CAGRA builds a high-degree graph optimized for GPU memory bandwidth, achieving faster search than HNSW on GPU but with higher memory requirements - the graph itself can be 2–3x larger than an equivalent HNSW index.
For production deployments, Weaviate with cuVS on B200 is particularly interesting. B200's 192 GB HBM3e can hold a CAGRA index for a 100M+ vector dataset entirely in GPU memory, eliminating PCIe round-trips. In this configuration, we observed 8,500 QPS at p99 latency of 12ms on a single B200 - the fastest single-node vector search result in our benchmarks.
Performance Comparison Across Platforms
We benchmarked Milvus, Qdrant, and Weaviate on identical H100 and B200 single-node configurations using the DBPedia 10M dataset (768-dimensional vectors) with recall target >= 95%. All measurements are steady-state QPS after 30-minute warmup.
Weaviate cuVS results use CAGRA index with default parameters. Milvus uses IVF_PQ with GPU coarse quantizer. Qdrant uses GPU-accelerated HNSW. Results are for single-node, single-GPU configurations.
| Platform | QPS (H100) | p99 Latency (H100) | QPS (B200) | $/1M Queries |
|---|---|---|---|---|
| Milvus IVF_PQ (GPU) | 3,800 | 38ms | 6,200 | $1.42 |
| Qdrant HNSW (GPU) | 2,100 | 72ms | 3,500 | $2.58 |
| Weaviate CAGRA | 4,900 | 22ms | 8,500 | $1.04 |
| Milvus (CPU only) | 950 | 112ms | N/A | $4.80 |
Cost Analysis at Scale
The cost-per-query analysis reveals that GPU-accelerated vector search is significantly cheaper than CPU at scale, but only if query volume is high enough to amortize GPU idle time. At 1M queries/month, a CPU-only Milvus cluster costs roughly $480/month in hardware at Lambda pricing ($2.80/GPU-hr equivalent for CPU). The GPU cluster costs roughly $620/month for 1M queries - more expensive.
The crossover is at roughly 10M queries/month. At 50M queries/month, GPU-accelerated search is 3.2x cheaper per query. B200 configurations cross over even faster (around 5M queries/month) because of the higher throughput per dollar. For production RAG pipelines serving 100M+ queries/month, GPU-accelerated vector search is not an optimization - it is a requirement to keep within latency SLOs.
When GPU Vector Search Makes Sense
GPU acceleration for vector search is a threshold game. Below 5M queries/month, the GPU idle cost dominates and CPU-only clusters are more economical. Between 5M–50M queries/month, GPU acceleration reduces latency but may not improve dollar efficiency depending on your index size and recall requirements. Above 50M queries/month, GPU acceleration is strictly better on both latency and cost.
The wildcard is B200 and B300 with larger HBM pools. A B200 can hold a 100M+ vector index entirely in GPU memory, eliminating PCIe transfers entirely. If your vector dataset fits in the GPU memory of a single node, GPU-accelerated search becomes both faster and cheaper than CPU at any query volume - even at 1M queries/month, the cost is comparable while latency is 4–6x better.
