All essays
TechnicalDEEP DIVEFEB 2026

GPU Inference with Vision Models: CLIP, SigLIP, ViT Serving Architecture

Technical deep dive on serving vision models for GPU inference: CLIP, SigLIP, ViT, and multimodal architectures. Vision encoder memory, batch processing, embedding caching, and vision-language fusion patterns for production.

01

THE VISION ENCODER LANDSCAPE IN 2026

Vision models for GPU inference span three primary architectures in 2026. CLIP (Contrastive Language-Image Pretraining) uses a dual-encoder architecture with a ViT-based image encoder and a transformer text encoder, producing aligned embedding spaces for image-text similarity. SigLTP (Sigmoid Loss for Language-Image Pretraining) improves on CLIP with a pairwise sigmoid loss that enables training on larger batch sizes and achieves 2-4% better zero-shot classification accuracy at equivalent model size. ViT (Vision Transformer) standalone models handle image classification, object detection, and segmentation without the text alignment component. Each architecture has distinct GPU inference characteristics that affect serving decisions.

The inference cost centers on the vision encoder, which processes image patches through transformer layers. A ViT-L/14 model (used in CLIP) at 224x224 resolution divides the image into 256 patches (14x14 each), each projected through a linear embedding and then processed through 24 transformer layers with 16 attention heads. The total compute is roughly 1.2 TFLOPs per image on H100, completing in 2-4ms at batch size 1. Higher-resolution variants (ViT-H/14 at 378x378, used in SigLIP) process 729 patches with 32 transformer layers at 3.8 TFLOPs per image, taking 6-10ms. For multimodal models that combine vision and language (LLaVA, Qwen-VL), the vision encoder runs once and the text decoder runs separately, meaning the vision cost is amortized across multiple text generation steps.

ModelParamsFLOPs per ImageLatency BS 1 (H100)
CLIP ViT-B/3286M0.3 TFLOPs1-2ms
CLIP ViT-L/14304M1.2 TFLOPs2-4ms
SigLIP ViT-L/16307M1.5 TFLOPs3-5ms
SigLIP ViT-H/14632M3.8 TFLOPs6-10ms
ViT-H (standalone)632M4.2 TFLOPs7-12ms
ViT-G/14 (EVA-02)1.0B6.8 TFLOPs12-18ms
02

IMAGE BATCH PROCESSING STRATEGIES

Vision model inference benefits significantly from batching because the vision encoder is compute-bound (dominated by matrix multiplications in the transformer layers) rather than memory-bandwidth bound like LLM decoders. On H100, CLIP ViT-L/14 achieves 250 images/second at BS 1, but scales to 4,200 images/second at BS 256 - a 16.8x throughput increase from 256x batching. This near-linear scaling makes vision model serving highly efficient for batched workloads like embedding extraction for vector databases, content moderation pipelines, and large-scale image classification.

The batch scaling does have limits. At BS 512+ on H100, GPU memory becomes the constraint. CLIP ViT-L/14 at BS 512 consumes approximately 18 GB of GPU memory for activations (intermediate states between transformer layers). Combined with model weights (0.6 GB at FP16), this fits comfortably within H100's 80 GB. Higher-resolution models like SigLIP ViT-H/14 at BS 512 consume 52 GB of activation memory, requiring careful memory budgeting. The optimal batch size for throughput-per-dollar on H100 is BS 128-256 for ViT-L models and BS 64-128 for ViT-H models. Beyond these ranges, diminishing returns from GPU compute saturation and increasing memory pressure make larger batches uneconomical.

03

IMAGE EMBEDDING CACHE FOR REPETITIVE PROCESSING

In production image pipelines, the same images are often processed multiple times. A content moderation pipeline, for example, may run CLIP inference on the same image for NSFW detection, brand detection, and aesthetic scoring. Without caching, each pipeline stage re-encodes the image, wasting GPU compute. An image embedding cache keyed on image hash (SHA256 of the decoded pixels) eliminates redundant encoding. The cache stores the vision encoder's output embedding (typically 512-1024 float32 values = 2-8 KB per image), reducing GPU cost for repeated processing by 60-90% depending on the reuse rate.

The cache can be implemented as a local LRU cache on each GPU instance (fastest, limited to GPU memory budget) or a distributed Redis cache (shared across instances, adds 1-3ms network latency). For content moderation at scale processing 1 million images per day with an average reuse factor of 3 (each image checked by 3 different classifiers), the embedding cache reduces daily GPU hours from 12 to 4 on a CLIP ViT-L serving pipeline - a 66% reduction. The cache storage cost is negligible: 1 million cached embeddings at 4 KB each is 4 GB, easily handled by a Redis cluster with 1 GB of memory overhead.

Cache StrategyHit RateGPU Time Saved
No cache0%0%
Per-GPU LRU cache40-55%35-45%
Distributed Redis cache55-70%50-65%
Local + distributed hybrid65-80%60-75%
Embedding + metadata cache75-90%70-85%
04

VISION-LANGUAGE FUSION SERVING PATTERNS

Multimodal models (LLaVA-NeXT, Qwen2-VL, GPT-4o-style architectures) combine vision encoding with LLM decoding. The serving architecture must handle the heterogeneous compute profile: the vision encoder is compute-bound and batch-friendly, while the decoder is memory-bandwidth bound and benefits from continuous batching. The standard pattern is a two-stage pipeline: stage 1 runs the vision encoder to produce visual embeddings, stage 2 concatenates these with text embeddings and runs the LLM decoder. Decoupling these stages allows independent scaling and cost optimization.

The vision encoding stage can run on a smaller GPU class (L40S at 16 GB) since it only needs memory for the vision model weights plus activations, typically 4-8 GB total. The decoder stage requires H100-class GPUs with large memory. Decoupled serving adds the overhead of transferring visual embeddings from the vision GPU to the decoder GPU, typically 1-3ms for 576 patch embeddings at 4,096 dimensions each (9 MB per image). For cost optimization on ClusterBid, the split architecture reduces per-request inference cost by 30-40% compared to running the full pipeline on H100, because the vision encoder runs on cheaper L40S GPUs at $0.35/hr versus $1.15/hr for H100. Total multimodal inference cost for a typical request with 1 image and 500 text tokens is approximately $0.0035 per request on the split architecture versus $0.0058 on unified H100.

05

PRODUCTION DEPLOYMENT PATTERNS FOR VISION INFERENCE

Vision inference deployments fall into two categories: embedding extraction (high-throughput, stateless) and multimodal chat (latency-sensitive, stateful). For embedding extraction, the deployment pattern is a stateless pool of H100 or L40S instances behind a load balancer, each running a vLLM or Triton serving the vision encoder model. Batching is aggressive (BS 128-256), and KV caching is irrelevant since there is no autoregressive decoding. The key metric is images per second per dollar, which for CLIP ViT-L on H100 is approximately 3,200 images/sec at $1.15/hr, or 10 million images per dollar.

For multimodal chat, the deployment follows the streaming architecture with the decoupled vision encoder pattern. The vision encoder GPU runs as a stateless preprocessor, while the decoder GPU maintains KV cache state for the conversation. Session affinity routes requests to the decoder GPU holding the conversation state, while the vision encoder is stateless and can be any available instance. This asymmetric architecture handles the fundamentally different compute profiles of vision encoding and text decoding. On ClusterBid, the recommended production configuration for multimodal serving at scale is a 2:1 ratio of L40S (vision encoding) to H100 (text decoding), achieving optimal cost-performance for mixed vision-text workloads.

Filed under
Vision Model InferenceCLIP GPU ServingSigLIP ArchitectureViT DeploymentMultimodal GPUVision-Language ModelsEmbedding Cache Vision