All essays
TechnicalDEEP DIVEFEB 2026

Inference with Embedding Models: text-embedding-3-large, voyage, bge Serving at Scale

Technical guide to serving embedding models at scale. text-embedding-3-large, voyage-3, and BGE-M3 comparison. Matryoshka embedding, pooling strategies, dimensional reduction, batch processing, and cost optimization.

01

EMBEDDING MODEL LANDSCAPE AND ARCHITECTURE

Embedding models convert text into dense vector representations. In 2026, the three leading families are: OpenAI's text-embedding-3-large (3,072 dimensions, transformer decoder architecture, 1.5B parameters), Voyage AI's voyage-3 (1,024 dimensions, proprietary transformer encoder, reported 750M parameters), and BGE-M3 from BAAI (1,024 dimensions, 567M parameters, supports dense and sparse retrieval, multilingual, up to 8K tokens). Each uses a different architecture for pooling token-level representations into a single text embedding, which directly affects inference characteristics and GPU requirements.

text-embedding-3-large is only available through OpenAI's API ($0.13 per 1M input tokens), while voyage-3 is available both through API ($0.12 per 1M tokens) and as a self-hosted model. BGE-M3 is fully open-source and self-hostable, making it the dominant choice for GPU clusters running embedding at scale. The self-hosting cost math: BGE-M3 on a single H100 handles approximately 12 million input tokens per hour at BS 256 with mean pooling. At $1.15/hr for H100 on ClusterBid, this equates to $0.096 per 1M tokens - cheaper than API pricing. At 100M tokens per day, self-hosting with BGE-M3 saves approximately $3,400 per month compared to API-based text-embedding-3-large.

ModelDimensionsCost per 1M Tokens
text-embedding-3-large3,072 (can reduce)$0.13 (API)
voyage-31,024$0.12 (API)
voyage-3 (self-host)1,024$0.06-0.08 (GPU)
BGE-M3 (dense + sparse)1,024$0.045-0.06 (GPU)
BGE-small-en-v1.5384$0.015-0.02 (GPU)
E5-mistral-7b4,096$0.18-0.25 (GPU)
02

POOLING STRATEGIES AND GPU COMPUTE COST

Pooling transforms per-token embeddings into a single text-level embedding. The three main strategies are: mean pooling (average token embeddings of the last hidden state), CLS token pooling (use only the first token's embedding), and attention-weighted pooling (learned weighted average). Mean pooling is the most common for open-source models and requires computing the full sequence through the encoder, then averaging over the sequence dimension - essentially free compute. CLS pooling is cheaper because it only needs the first token's representation, but most embedding models use mean pooling or attention-weighted variants.

The GPU cost of embedding generation is dominated by the encoder forward pass, not pooling. BGE-M3 at sequence length 512 processes approximately 4,100 tokens per second per H100 at BS 1, and 36,000 tokens/sec at BS 256. The encoder compute is proportional to sequence length and batch size. Mean pooling adds approximately 0.2-0.5% additional compute because it's a simple element-wise mean of the hidden states. The practical implication is that the sequence length per input is the dominant cost factor: a 512-token input costs 8x more than a 64-token input to embed, not 8x more per token but roughly 8x more compute because transformer self-attention is O(n^2) in sequence length. For RAG pipelines, keeping input texts under 256 tokens per segment significantly reduces embedding cost.

03

MATRYOSHKA EMBEDDINGS AND DIMENSIONALITY REDUCTION

Matryoshka Representation Learning (MRL) trains embedding models to produce representations where the first N dimensions form a valid embedding for any N. This allows downstream applications to use fewer dimensions without retraining. text-embedding-3-large supports matryoshka truncation: you can request embeddings at 256, 512, 1,024, or the full 3,072 dimensions. BGE-M3 also supports matryoshka-style truncation at 256, 512, 768, and 1,024 dimensions. The compute savings come entirely from the vector database index: reducing dimensions from 3,072 to 256 cuts vector storage costs by 8x and ANN search cost by 4-6x, with only 1-3% accuracy loss on retrieval benchmarks.

The GPU cost of generating embeddings is the same regardless of output dimension in matryoshka models, because the full model must compute all dimensions before truncation. The savings are downstream in storage and search. For large-scale RAG systems with 100 million vectors, the storage cost math: 3,072 dimensions at FP32 = 12 KB per vector = 1.2 TB total; 256 dimensions = 1 KB per vector = 100 GB total. Using 256 dimensions from text-embedding-3-large reduces vector database cost from approximately $600/month to $80/month on a managed Milvus cluster, while maintaining 95%+ of the retrieval accuracy of 3,072 dimensions. The optimal dimension is determined by benchmarking your specific retrieval task at 256, 512, 1,024, and 3,072 dimensions - most RAG workloads see diminishing returns beyond 512.

Output DimensionsVector Storage (100M)Retrieval Accuracy vs Full
3,072 (full)1.2 TB100% (baseline)
1,024400 GB98-99%
512200 GB96-98%
256100 GB94-97%
12850 GB89-94%
04

BATCH PROCESSING FOR EMBEDDING THROUGHPUT

Embedding models are highly batchable because the encoder transformer is compute-bound and processes all batch elements in parallel. BGE-M3 on H100 achieves 4,100 tokens/sec at BS 1 (approximately 8 texts of 512 tokens per second), but scales to 36,000 tokens/sec at BS 256 (70 texts per second). The scaling efficiency from BS 1 to BS 256 is approximately 8.8x - less than linear because the GPU compute capacity saturates and memory bandwidth for weight reads becomes the limiting factor. The sweet spot for cost-efficiency is BS 64-128 on H100, where throughput-per-watt and throughput-per-dollar peak.

The key throughput optimization is packing: grouping short texts together in a single batch to maximize tokens per GPU iteration. A batch of 128 texts averaging 128 tokens each achieves the same throughput as 8 texts of 2,048 tokens, but processes 16x more texts per second. The vLLM embedding server and TEI (Text Embeddings Inference) automatically pack batches by padding to the longest sequence in each batch, but variable-length packing across batches is imperfect. For maximum throughput, sort incoming texts by length and batch similar-length texts together, reducing padding waste from 30-50% to 5-15%. This sort-and-batch strategy improves effective throughput by 25-40% on H100, reducing cost per embedded text by the same margin.

Batch StrategyTokens/sec (H100)Texts/sec (avg 256 tok)
BS 1 (no batching)4,10016
BS 1612,50049
BS 6425,00098
BS 12832,000125
BS 25636,000141
BS 256 + sort-pack44,000172
05

PRODUCTION EMBEDDING SERVING ARCHITECTURE

The production embedding serving architecture diverges from LLM serving in three important ways. First, there is no KV cache or session state: embedding requests are stateless and can be routed to any GPU. Second, the latency requirement is typically looser: 200-500ms is acceptable for embedding in RAG pipelines, versus 50-200ms for chat. Third, batching is the dominant optimization lever, whereas LLM serving optimizes for continuous batching with variable generation lengths. The optimal architecture is a stateless GPU pool running vLLM in embedding mode or TEI, with a request buffer that accumulates embeddings into optimal batch sizes before dispatching to GPUs.

For high-scale deployments processing 10M+ texts per day, the architecture uses an intake buffer (Redis or Kafka) that collects embedding requests and releases them to the GPU pool in sorted batches of 64-128. The GPU pool runs on H100 or L40S (L40S at $0.35/hr provides 70% of H100 embedding throughput at 30% of the cost, making it the most cost-effective embedding GPU). The buffer-to-GPU dispatch achieves 95%+ GPU utilization by maintaining a steady queue depth. On ClusterBid, the recommended configuration is a pool of 4-8 L40S GPUs for embedding serving, achieving throughput of 50-100M tokens per hour at $0.02-0.03 per 1M tokens - 4-6x cheaper than API-based embedding services. For teams already running LLM inference on H100s, the same GPUs can serve embedding requests during off-peak hours without additional GPU cost, making embedding essentially free when colocated with chat inference.

GPU ConfigurationTokens/HourCost per 1M Tokens
1x L40S (embeddings)8-12M$0.03-0.045
4x L40S (embeddings)35-50M$0.02-0.03
1x H100 (embeddings)35-45M$0.025-0.035
8x H100 (embeddings)280-360M$0.02-0.03
API: text-embedding-3Unlimited$0.13
Self-hosted (colocated)Depends on headroom$0.00-0.01
Filed under
Embedding Model InferenceText Embedding ServingMatryoshka EmbeddingBGE-M3 GPUVoyage AI ServingBatch EmbeddingVector Database Pipeline