All essays
TechnicalDEEP DIVEFEB 2026

Embedding Model GPU Infrastructure: Text, Image, and Multi-Modal Embedding Serving at Scale

Production GPU infrastructure for embedding models: text-embedding-3-large, BGE-M3, Voyage-3, SigLIP, and multimodal CLIP serving. Throughput benchmarks, batch encoding strategies, VRAM requirements, and cost per 1M embeddings on L40S, A100, and H100 for search and RAG pipelines.

01

EMBEDDING MODEL COMPUTE PROFILES

Embedding models are encoder-only Transformers or dual-encoder architectures that process input into a single dense vector. Unlike decoder LLMs, they run one forward pass per input without autoregressive generation. This single-pass characteristic makes embedding serving a throughput-oriented, memory-bound workload dominated by the encoder's attention and pooler layers. Text-embedding-3-large (unknown but estimated 300M params, 1,536 dim) runs one forward pass of 8.5 ms on H100 for 512-token input at batch size 1, producing one 1,536-dim vector (6 KB). The compute-to-output ratio is extremely high: 8.5 ms of H100 compute for 6 KB of output. Batch processing is essential to amortize the model load cost.

BGE-M3 (567M params, multi-lingual, supports 8,192 token input) uses a hybrid architecture: birectional encoder with MCLS pooling and optional ColBERT-style late interaction for fine-grained retrieval. At 8,192-token input on H100, a single forward pass takes 340 ms at BS=1 due to the quadratic attention over 8,192 tokens. On A100, the same forward pass takes 620 ms. The memory for a single BGE-M3 forward pass: 1.1 GB weights + 1.3 GB attention activation maps at 8K context = 2.4 GB. At max batch size 128 on H100, activation memory reaches 128 * 1.3 GB = 166 GB, exceeding H100's 80 GB. The effective max batch is 48, where each pass takes 14.2 seconds total and produces 48 embeddings. Throughput at BS=48: 3.4 embeddings per second for 8K inputs, versus 3,200 embeddings per second for 128-token inputs at BS=256.

ModelParamsDimMax TokensBS=1 Latency (H100)Max BS (H100 80GB)Embs/sec H100 Max BS
text-embedding-3-small~150M5128,1913.5 ms (512 tok)1,02442,000
text-embedding-3-large~300M1,5368,1918.5 ms (512 tok)51224,000
BGE-M3567M1,0248,19212 ms (512 tok)51221,000
BGE-M3 (long)567M1,0248,192340 ms (8K tok)48140
Voyage-3~250M1,02432,00015 ms (512 tok)1284,800
text-embedding-3-large-3k~300M1,5363,0724.2 ms (512 tok)51228,000
02

GPU COMPARISON: THROUGHPUT AND COST PER 1M EMBEDDINGS

The cost per million embeddings varies significantly by GPU choice and embedding model. For text-embedding-3-large at 512-token average input with BS=256 on H100 ($2.50/hr), throughput is 24,000 embeddings/sec, costing $0.029 per 1M embeddings. On A100 80GB PCIe ($1.75/hr) at BS=128, throughput is 11,000 embeddings/sec, costing $0.044 per 1M. On L40S ($1.10/hr) at BS=64, throughput is 6,200 embeddings/sec, costing $0.049 per 1M. On A10 ($0.60/hr) at BS=16, throughput is 1,500 embeddings/sec, costing $0.111 per 1M.

For BGE-M3 at 512-token average input, the rankings shift. The larger model (567M) is more bandwidth-bound, and H100's 3.35 TB/s bandwidth gives it a larger advantage: H100 at 21,000 emb/s ($0.033/1M), A100 at 8,400 emb/s ($0.058/1M), L40S at 5,200 emb/s ($0.059/1M), A10 at 1,100 emb/s ($0.152/1M). For long-context BGE-M3 at 8K tokens, the bandwidth advantage is even more pronounced: H100 at 140 emb/s ($4.96/1M), A100 at 52 emb/s ($9.35/1M). The L40S cannot run BGE-M3 at 8K tokens at BS > 1 because the attention maps (1.3 GB per input at 8K) exceed its 48 GB VRAM at BS=12. For long-context embedding serving, H100 is the only viable GPU.

GPUEmb/s text-3-large (512t)$/1M embEmb/s BGE-M3 (512t)$/1M embEmb/s BGE-M3 (8Kt)$/1M emb
A10 24GB1,500$0.1111,100$0.152Not feasibleN/A
L40S 48GB6,200$0.0495,200$0.059Not feasibleN/A
A100 80GB PCIe11,000$0.0448,400$0.05852$9.35
H100 SXM 80GB24,000$0.02921,000$0.033140$4.96
H200 141GB32,000$0.03028,000$0.034210$4.50
03

MULTI-MODAL AND IMAGE EMBEDDING SERVING

Image embedding models require significantly more compute per vector than text embedding. SigLIP-SO400M (400M params) converts a 448px image to a single 1,152-dim embedding in 1.5 ms on H100 at BS=1, producing 0.5 MB of output for the full batch of feature maps before pooling. At BS=128, throughput reaches 42,000 embeddings/second on H100, costing $0.016 per 1M embeddings. CLIP ViT-L/14 (428M params) at 224px: 0.4 ms per image at BS=1, 85,000 emb/s at BS=256 on H100, $0.008 per 1M. OpenCLIP ViT-H/14 (632M params) at 224px: 1.1 ms per image, 38,000 emb/s at BS=128, $0.018 per 1M.

The cost crossover: for pure text embedding, L40S at $0.049/1M is competitive with H100 at $0.029/1M for moderate-volume pipelines. For image embedding, the throughput-per-dollar advantage of H100 is larger due to faster batch processing: H100 achieves 42,000 SigLIP emb/s ($0.016/1M) versus L40S 14,000 emb/s ($0.022/1M). For multi-modal embedding pipelines that produce both text and image embeddings from the same model (like CLIP), the choice depends on the text-to-image ratio. At 80% text + 20% image, the blended cost is $0.026/1M on H100 versus $0.044/1M on L40S. At 95% text + 5% image, L40S at $0.047/1M is close to H100 at $0.028/1M, making L40S viable for text-heavy pipelines at half the GPU cost.

04

CLUSTER SIZING FOR PRODUCTION EMBEDDING PIPELINES

Production embedding pipelines at search and RAG scale process 100M-1B documents for initial indexing plus incremental updates. Document indexing is throughput-oriented and latency-tolerant: a batch pipeline can process documents over hours using large batch sizes and CPU preprocessing. For a 500M document index with text-embedding-3-large at 512 tokens average, each document takes 41.7 microseconds on H100 (24,000 emb/s). Total compute time: 500M * 41.7 us = 20,833 seconds = 5.8 hours on a single H100. With 8 H100 GPUs, index time drops to 43 minutes. At $2.50/hr per GPU, total indexing cost is $116 on 8 GPUs or $14.50 on 1 GPU.

Query-time embedding serving has different constraints: sub-50ms per query for real-time search, latency-sensitive. A single H100 with text-embedding-3-large at BS=1 processes 118 queries per second (8.5 ms each). A RAG pipeline running 50 QPS requires 1 H100 GPU for embeddings plus 4 H100s for LLM inference and vector search. The embedding GPU is rarely the bottleneck: a single H100 handles 10M queries/day. The real infrastructure cost is the LLM inference cluster, not the embedding tier. For multi-modal search serving 100 QPS with 1 image + 64 text tokens per query, SigLIP on 1 H100 handles 680 images/sec at BS=4, well above the 100 QPS requirement, with the embedding GPU costing only $0.007 per 1K queries.

Filed under
Embedding Model GPUText Embedding ServingBGE-M3 GPUCLIP GPU InferenceMultimodal Embedding GPURAG Embedding InfrastructureVector Search GPU