EMBEDDING MODEL COMPUTE PROFILES
Embedding models are encoder-only Transformers or dual-encoder architectures that process input into a single dense vector. Unlike decoder LLMs, they run one forward pass per input without autoregressive generation. This single-pass characteristic makes embedding serving a throughput-oriented, memory-bound workload dominated by the encoder's attention and pooler layers. Text-embedding-3-large (unknown but estimated 300M params, 1,536 dim) runs one forward pass of 8.5 ms on H100 for 512-token input at batch size 1, producing one 1,536-dim vector (6 KB). The compute-to-output ratio is extremely high: 8.5 ms of H100 compute for 6 KB of output. Batch processing is essential to amortize the model load cost.
BGE-M3 (567M params, multi-lingual, supports 8,192 token input) uses a hybrid architecture: birectional encoder with MCLS pooling and optional ColBERT-style late interaction for fine-grained retrieval. At 8,192-token input on H100, a single forward pass takes 340 ms at BS=1 due to the quadratic attention over 8,192 tokens. On A100, the same forward pass takes 620 ms. The memory for a single BGE-M3 forward pass: 1.1 GB weights + 1.3 GB attention activation maps at 8K context = 2.4 GB. At max batch size 128 on H100, activation memory reaches 128 * 1.3 GB = 166 GB, exceeding H100's 80 GB. The effective max batch is 48, where each pass takes 14.2 seconds total and produces 48 embeddings. Throughput at BS=48: 3.4 embeddings per second for 8K inputs, versus 3,200 embeddings per second for 128-token inputs at BS=256.
| Model | Params | Dim | Max Tokens | BS=1 Latency (H100) | Max BS (H100 80GB) | Embs/sec H100 Max BS |
|---|---|---|---|---|---|---|
| text-embedding-3-small | ~150M | 512 | 8,191 | 3.5 ms (512 tok) | 1,024 | 42,000 |
| text-embedding-3-large | ~300M | 1,536 | 8,191 | 8.5 ms (512 tok) | 512 | 24,000 |
| BGE-M3 | 567M | 1,024 | 8,192 | 12 ms (512 tok) | 512 | 21,000 |
| BGE-M3 (long) | 567M | 1,024 | 8,192 | 340 ms (8K tok) | 48 | 140 |
| Voyage-3 | ~250M | 1,024 | 32,000 | 15 ms (512 tok) | 128 | 4,800 |
| text-embedding-3-large-3k | ~300M | 1,536 | 3,072 | 4.2 ms (512 tok) | 512 | 28,000 |
GPU COMPARISON: THROUGHPUT AND COST PER 1M EMBEDDINGS
The cost per million embeddings varies significantly by GPU choice and embedding model. For text-embedding-3-large at 512-token average input with BS=256 on H100 ($2.50/hr), throughput is 24,000 embeddings/sec, costing $0.029 per 1M embeddings. On A100 80GB PCIe ($1.75/hr) at BS=128, throughput is 11,000 embeddings/sec, costing $0.044 per 1M. On L40S ($1.10/hr) at BS=64, throughput is 6,200 embeddings/sec, costing $0.049 per 1M. On A10 ($0.60/hr) at BS=16, throughput is 1,500 embeddings/sec, costing $0.111 per 1M.
For BGE-M3 at 512-token average input, the rankings shift. The larger model (567M) is more bandwidth-bound, and H100's 3.35 TB/s bandwidth gives it a larger advantage: H100 at 21,000 emb/s ($0.033/1M), A100 at 8,400 emb/s ($0.058/1M), L40S at 5,200 emb/s ($0.059/1M), A10 at 1,100 emb/s ($0.152/1M). For long-context BGE-M3 at 8K tokens, the bandwidth advantage is even more pronounced: H100 at 140 emb/s ($4.96/1M), A100 at 52 emb/s ($9.35/1M). The L40S cannot run BGE-M3 at 8K tokens at BS > 1 because the attention maps (1.3 GB per input at 8K) exceed its 48 GB VRAM at BS=12. For long-context embedding serving, H100 is the only viable GPU.
| GPU | Emb/s text-3-large (512t) | $/1M emb | Emb/s BGE-M3 (512t) | $/1M emb | Emb/s BGE-M3 (8Kt) | $/1M emb |
|---|---|---|---|---|---|---|
| A10 24GB | 1,500 | $0.111 | 1,100 | $0.152 | Not feasible | N/A |
| L40S 48GB | 6,200 | $0.049 | 5,200 | $0.059 | Not feasible | N/A |
| A100 80GB PCIe | 11,000 | $0.044 | 8,400 | $0.058 | 52 | $9.35 |
| H100 SXM 80GB | 24,000 | $0.029 | 21,000 | $0.033 | 140 | $4.96 |
| H200 141GB | 32,000 | $0.030 | 28,000 | $0.034 | 210 | $4.50 |
MULTI-MODAL AND IMAGE EMBEDDING SERVING
Image embedding models require significantly more compute per vector than text embedding. SigLIP-SO400M (400M params) converts a 448px image to a single 1,152-dim embedding in 1.5 ms on H100 at BS=1, producing 0.5 MB of output for the full batch of feature maps before pooling. At BS=128, throughput reaches 42,000 embeddings/second on H100, costing $0.016 per 1M embeddings. CLIP ViT-L/14 (428M params) at 224px: 0.4 ms per image at BS=1, 85,000 emb/s at BS=256 on H100, $0.008 per 1M. OpenCLIP ViT-H/14 (632M params) at 224px: 1.1 ms per image, 38,000 emb/s at BS=128, $0.018 per 1M.
The cost crossover: for pure text embedding, L40S at $0.049/1M is competitive with H100 at $0.029/1M for moderate-volume pipelines. For image embedding, the throughput-per-dollar advantage of H100 is larger due to faster batch processing: H100 achieves 42,000 SigLIP emb/s ($0.016/1M) versus L40S 14,000 emb/s ($0.022/1M). For multi-modal embedding pipelines that produce both text and image embeddings from the same model (like CLIP), the choice depends on the text-to-image ratio. At 80% text + 20% image, the blended cost is $0.026/1M on H100 versus $0.044/1M on L40S. At 95% text + 5% image, L40S at $0.047/1M is close to H100 at $0.028/1M, making L40S viable for text-heavy pipelines at half the GPU cost.
CLUSTER SIZING FOR PRODUCTION EMBEDDING PIPELINES
Production embedding pipelines at search and RAG scale process 100M-1B documents for initial indexing plus incremental updates. Document indexing is throughput-oriented and latency-tolerant: a batch pipeline can process documents over hours using large batch sizes and CPU preprocessing. For a 500M document index with text-embedding-3-large at 512 tokens average, each document takes 41.7 microseconds on H100 (24,000 emb/s). Total compute time: 500M * 41.7 us = 20,833 seconds = 5.8 hours on a single H100. With 8 H100 GPUs, index time drops to 43 minutes. At $2.50/hr per GPU, total indexing cost is $116 on 8 GPUs or $14.50 on 1 GPU.
Query-time embedding serving has different constraints: sub-50ms per query for real-time search, latency-sensitive. A single H100 with text-embedding-3-large at BS=1 processes 118 queries per second (8.5 ms each). A RAG pipeline running 50 QPS requires 1 H100 GPU for embeddings plus 4 H100s for LLM inference and vector search. The embedding GPU is rarely the bottleneck: a single H100 handles 10M queries/day. The real infrastructure cost is the LLM inference cluster, not the embedding tier. For multi-modal search serving 100 QPS with 1 image + 64 text tokens per query, SigLIP on 1 H100 handles 680 images/sec at BS=4, well above the 100 QPS requirement, with the embedding GPU costing only $0.007 per 1K queries.
