GPU LANDSCAPE FOR SMALL MODEL SERVING
Small language models (1-7B) present a unique GPU cost equation: memory bandwidth dominates over compute capacity because the attention mechanism processes a tiny parameter set, and the memory bottleneck is moving weights through HBM into the compute units. A Llama 3.2 1B model (0.5 GB at FP16) on H100 with 3.35 TB/s bandwidth should theoretically achieve 6,700 tok/s memory-bound throughput. Real throughput at batch size 1 is 4,200 tok/s, limited by kernel launch overhead and attention compute. The memory bandwidth utilization for SLMs is only 15-25% versus 60-70% for 70B models, because the compute-to-weight ratio is too high for the weights to stay in cache.
The four GPUs in this comparison span a 5x cost range and 7x throughput range. L40S ($1.10/hr) with 48 GB VRAM and 1.35 TB/s bandwidth. A10 ($0.60/hr) with 24 GB, 600 GB/s. A100 80GB PCIe ($1.75/hr) with 2.0 TB/s. H100 SXM ($2.50/hr) with 3.35 TB/s. For SLMs, the GPU choice depends on whether the deployment is single-user latency-optimized (batch size 1), high-throughput batched, or multi-model shared. Each use case selects a different GPU on the price-performance curve.
| GPU | VRAM | HBM BW | $/hr | Llama 3.2 1B tok/s BS=1 | Qwen 2.5 7B tok/s BS=1 |
|---|---|---|---|---|---|
| L40S | 48 GB | 1.35 TB/s | $1.10 | 1,850 | 340 |
| A10 | 24 GB | 600 GB/s | $0.60 | 820 | 140 |
| A100 80GB PCIe | 80 GB | 2.0 TB/s | $1.75 | 2,900 | 540 |
| H100 SXM 80GB | 80 GB | 3.35 TB/s | $2.50 | 4,200 | 810 |
| H200 SXM 141GB | 141 GB | 4.8 TB/s | $3.40 | 5,800 | 1,080 |
BATCH INFERENCE THROUGHPUT COMPARISON
Batch inference changes the GPU ranking. At batch size 64 for Llama 3.2 3B, H100 achieves 42,000 tok/s, A100 achieves 26,000 tok/s, L40S achieves 18,500 tok/s, and A10 achieves 7,200 tok/s. Cost per 1M tokens: H100 $0.017, A100 $0.019, L40S $0.017, A10 $0.023. L40S matches H100 on cost per token at high batch sizes despite 2.2x lower throughput, because its $1.10/hr rate offsets the throughput gap. At batch size 1 latency-critical serving, H100 is the cost leader at $0.17 per 1M tokens for Llama 3.2 1B, with the higher throughput overcoming the higher hourly cost.
For Qwen 2.5 7B, the VRAM constraint on A10 becomes binding: 7B at FP16 is 14 GB + 4 GB KV cache at 8K context = 18 GB, leaving only 6 GB for batch accumulation. Max batch size on A10 is 8, achieving 720 tok/s at $0.23 per 1M tokens. L40S with 48 GB runs Qwen 2.5 7B at BS=64 achieving 9,800 tok/s at $0.031 per 1M tokens. H100 at BS=128 achieves 32,000 tok/s at $0.022 per 1M tokens. For 7B models, L40S is the cost leader by 5-10% over H100 on a pure cost-per-token basis, while H100 wins on throughput density (tokens per GPU per second).
| Model + GPU | Max BS on GPU | Peak tok/s | $/hr | $ per 1M tok (peak) | Req/sec for 1M tok/day |
|---|---|---|---|---|---|
| Llama 3.2 1B, A10 | 512 | 35,000 | $0.60 | $0.005 | 0.3 |
| Llama 3.2 1B, L40S | 2,048 | 92,000 | $1.10 | $0.003 | 0.1 |
| Llama 3.2 1B, A100 | 4,096 | 168,000 | $1.75 | $0.003 | 0.06 |
| Llama 3.2 1B, H100 | 8,192 | 280,000 | $2.50 | $0.003 | 0.04 |
| Qwen 2.5 7B, A10 | 8 | 720 | $0.60 | $0.231 | 14 |
| Qwen 2.5 7B, L40S | 64 | 9,800 | $1.10 | $0.031 | 1.0 |
| Qwen 2.5 7B, A100 | 128 | 18,000 | $1.75 | $0.027 | 0.6 |
| Qwen 2.5 7B, H100 | 256 | 32,000 | $2.50 | $0.022 | 0.3 |
MULTI-MODEL SERVING AND GPU PACKING
SLMs are small enough that multiple models can be packed onto a single GPU using multi-LoRA serving or model multiplexing. A single L40S with 48 GB can load Llama 3.2 3B (6 GB FP16), Gemma 2 2B (4 GB), and Qwen 2.5 1.5B (3 GB) simultaneously, with 35 GB remaining for KV caches and batch accumulation. Using vLLM multi-LoRA, 10 fine-tuned adapters can be applied on top of the same base Llama 3.2 3B, each occupying 0.1-0.3 GB, with 46 GB remaining for batch processing across all adapters. This enables a single L40S to serve 10 fine-tuned models at 120 tok/s each with 32ms per-token latency.
The cost advantage of model packing is significant. Serving 10 fine-tuned 3B adapters individually on the cheapest GPU (A10 at $0.60/hr each) would cost $6.00/hr. Packing them on a single L40S at $1.10/hr reduces cost by 82%. The throughput impact is modest: per-adapter throughput drops from 240 tok/s on dedicated A10 to 120 tok/s on packed L40S, but the aggregate throughput increases from 240 to 1,200 tok/s. For workloads with low concurrency per adapter but many adapters total, GPU packing on L40S or A100 delivers 3-5x better cost efficiency than per-model GPU allocation.
TOTAL COST OF OWNERSHIP SCENARIOS
Three production scenarios illustrate the GPU choice dynamics. Scenario 1: Chat assistant serving 10M tokens daily with Llama 3.2 3B, low latency requirement (sub-200ms). H100 handles this on 1 GPU at $60/month GPU cost. L40S would require 1 GPU at $26/month but delivers 340ms P99 latency. A10 delivers 850ms P99 latency at $14/month. For the latency-sensitive use case, L40S is the optimal choice meeting latency targets at 43% of H100 cost.
Scenario 2: High-throughput batch processing of 500M tokens daily with Qwen 2.5 7B. H100 cluster (4 GPUs) processes this in 4.3 hours/day at $68/day. L40S cluster (8 GPUs) processes in 12 hours/day at $106/day. A100 cluster (6 GPUs) processes in 8.5 hours/day at $119/day. For high-throughput batch, H100 wins on absolute cost due to higher throughput per GPU. Scenario 3: Multi-adapter serving with 50 fine-tuned 1B models, each serving 1M tokens/day. A single L40S packs all 50 adapters at $26/day. H100 packs all 50 adapters at $60/day with headroom for higher latency SLAs. L40S wins for multi-adapter workloads with a 57% cost reduction.
