VIT INFERENCE COMPUTE PROFILE
Vision Transformers process images by dividing them into fixed-size patches (typically 16x16 pixels) and projecting each patch into a token embedding. A ViT-H (632M parameters) processing a 224x224 image produces 196 tokens from 14x14 patches. The compute profile is compute-bound, not memory-bound: the attention operations over 196 tokens at 1,280 hidden dimension are small enough that the tensor core utilization reaches only 18-25% on H100. The bottleneck is patch embedding matrix multiply at the input and the classification head at the output. For ViT-L (304M parameters), a single-image forward pass on H100 takes 1.2 ms, while ViT-H takes 3.8 ms. The memory footprint is small: ViT-H at FP16 requires 1.3 GB for weights, plus 0.2 GB for activations at 224px resolution.
The small per-image compute means GPU throughput is primarily limited by batch size rather than memory. On an H100 SXM with 80 GB, ViT-L can batch 4,096 images before VRAM constraints appear (5.3 GB total), achieving 260,000 images per second at the model level. The limiting factor becomes CPU-side preprocessing: image decoding (JPEG/PNG), resize, normalize, and patchify. A single H100 running ViT-L at full throughput requires 32 CPU cores just to keep up with image preprocessing using TorchVision's dataloader, making the data pipeline the actual bottleneck in production ViT serving.
| Model | Params | VRAM (FP16) | Latency/Image BS=1 | Max BS on H100 | Max Images/s on H100 |
|---|---|---|---|---|---|
| ViT-B/16 | 86M | 0.2 GB | 0.4 ms | 16,384 | 820,000 |
| ViT-L/16 | 304M | 0.6 GB | 1.2 ms | 8,192 | 450,000 |
| ViT-H/16 | 632M | 1.3 GB | 3.8 ms | 4,096 | 260,000 |
| ViT-G/14 | 1.8B | 3.6 GB | 9.2 ms | 2,048 | 95,000 |
| SigLIP-SO400M | 400M | 0.8 GB | 1.5 ms | 8,192 | 380,000 |
SIGLIP AND OPENCLIP IMAGE-TOWER SERVING
SigLIP replaces the standard softmax-based contrastive loss with sigmoid loss, removing the need for large batch sizes during training, but inference is identical to ViT: patch embedding plus transformer layers. The common deployment pattern pairs SigLIP-SO400M (400M param image tower) with a separate text encoder, typically a 300M-parameter Transformer, for image-text embedding. On a single L40S, SigLIP-SO400M serves 42,000 images per second at batch size 2,048 with 7.8 ms latency per batch, drawing 7.2 GB VRAM. The same model on H100 achieves 380,000 images per second. L40S's advantage is the 48 GB VRAM allows larger batches for very-high-resolution images: 1,024px images at BS 256 require 18 GB on L40S versus 22 GB on H100 (similar), but L40S costs $1.10/hr versus H100 at $2.50/hr.
OpenCLIP's ViT-H-14 at 448px resolution (33.6 tokens from 32x32 patches) serves as a common production embedding model. At batch size 512 on A100 (80 GB PCIe), throughput is 94,000 images/second at $1.75/hr, cost per 1M images of $0.005. On H100, same batch, 210,000 images/second at $0.003 per 1M images. The L40S cost advantage ($1.10/hr) at batch size 1,024 yields 74,000 images/second, $0.004 per 1M images. For image embedding workloads, L40S offers the best price-performance at $0.004 per 1M images versus H100's $0.003 and A100's $0.005, making it the most cost-effective GPU for visual embedding pipelines that don't require maximum throughput.
SAM 2 AND DINOV2: SEGMENTATION AND FEATURE EXTRACTION
SAM 2 replaces SAM's heavyweight image encoder (ViT-H 632M) with a smaller Hiera backbone (83-680M variants) combined with a mask decoder and memory attention module for video. The image encoder processes at 1,024x1,024 resolution, producing tokens from 16x16 patches. Each image takes 45 ms on H100 for the encoder alone, plus 3-6 ms per prompt for the lightweight decoder. A typical production pipeline runs 4 encoder instances across 4 H100s with a shared decoder GPU, processing 12 images/second total. SAM 2's memory requirement is dominated by the Hiera encoder activations: at 1,024px, activations reach 2.5 GB per image for the base variant, limiting batch size to 16 on a single H100.
DINOv2's self-supervised ViT models are used for feature extraction at 518x518 resolution. The ViT-g/14 (1.1B params) uses register tokens and an output CLS token for global features. On A100, DINOv2-g processes 440 images/second at batch size 64, with 1.8 GB VRAM. The ViT-b/14 (86M) serves 2,900 images/second on the same GPU. For high-throughput feature extraction pipelines processing 10M+ images daily, DINOv2-b on L40S offers the best cost efficiency at $0.001 per 1K images, while DINOv2-g on H100 provides highest quality embeddings at $0.008 per 1K images.
| Model | GPU | Max BS | Images/sec | VRAM Used | Cost/1K Images |
|---|---|---|---|---|---|
| DINOv2-b/14 | L40S | 256 | 2,100 | 3.8 GB | $0.001 |
| DINOv2-b/14 | H100 | 512 | 5,800 | 4.2 GB | $0.001 |
| DINOv2-g/14 | A100 | 64 | 440 | 8.1 GB | $0.005 |
| DINOv2-g/14 | H100 | 128 | 980 | 9.4 GB | $0.004 |
| SAM 2 Hiera-L | H100 | 16 | 95 | 38 GB | $0.009 |
| SAM 2 Hiera-B | A100 | 8 | 35 | 32 GB | $0.014 |
PRODUCTION VIT SERVING: BATCHING AND LATENCY TRADEOFFS
ViT models benefit from unusual batching dynamics. Because the patch token sequence length is independent of batch size and remains short (196-256 tokens for standard resolutions), the attention computation scales linearly with batch size without the quadratic penalty seen in LLMs. This means ViT batch scaling is near-linear up to GPU memory limits: doubling batch size doubles throughput with less than 5% latency increase per image at H100 batch sizes up to 4,096. The practical throughput ceiling is 820,000 images/s for ViT-B and 260,000 for ViT-H on a single H100, limited by the patch embedding GEMM and classifier head rather than the transformer blocks.
For latency-critical applications like real-time video processing at 30 FPS, the batch size must be kept small. ViT-L at BS=1 on H100 achieves 1.2 ms per frame, supporting 833 FPS theoretical throughput. When preprocessing (decode + resize + normalize) adds 3-5 ms per frame on the same GPU, the effective frame rate drops to 200-250 FPS. Offloading preprocessing to CPU reduces GPU idle time: with 32 CPU cores preprocessing frames into pinned memory, a single H100 processes 720p video at 60 FPS with latency under 17 ms end-to-end. For offline batch processing, 4x H100 in a single node with TP=2 achieves 920,000 images/sec for ViT-H by reducing the attention GEMM memory pressure through sharding.
