THE LoRA COMPUTE GRAPH AND MEMORY FOOTPRINT
LoRA fine-tuning produces adapter weights that are orders of magnitude smaller than full model weights. A rank-16 LoRA adapter on a 70B model requires only 160 MB of additional parameters per layer. With 80 layers, the full LoRA stack is approximately 13 MB per adapter at FP16. This compact footprint is what makes multi-LoRA serving feasible: you can store thousands of adapters in host memory and swap them onto the GPU on demand. The base model weights remain fixed in GPU memory and are shared across all adapters.
The compute graph for LoRA inference modifies each linear layer: instead of the standard y = Wx, the forward pass becomes y = Wx + B(Ax) where A and B are the low-rank adapter matrices. This adds two small matrix multiplications per layer. For rank-16 adapters on a 70B model, the LoRA computation adds approximately 3-5% FLOPs overhead per forward pass compared to the base model alone. The overhead is proportional to adapter rank and decreases as a percentage of total compute for larger base models. At batch size 1, the LoRA overhead per token is roughly 0.3 ms on H100 for a 70B model.
| Model Size | LoRA Rank | Adapter Size (FP16) |
|---|---|---|
| 7B (32 layers) | Rank 16 | 0.5 MB |
| 7B (32 layers) | Rank 64 | 2.1 MB |
| 13B (40 layers) | Rank 16 | 1.0 MB |
| 70B (80 layers) | Rank 16 | 4.0 MB |
| 70B (80 layers) | Rank 128 | 32 MB |
| 405B (126 layers) | Rank 16 | 10 MB |
ADAPTER SCHEDULER ARCHITECTURE
vLLM implements multi-LoRA serving by maintaining a fixed-size GPU memory pool for adapter weights and swapping adapters in and out as requests arrive. When a request specifies a LoRA adapter, vLLM checks if the adapter is already resident in GPU memory. If not, it loads the adapter from host memory into a vacant slot, potentially evicting a least-recently-used adapter. The scheduler globally coordinates which adapters to keep resident based on request patterns, achieving optimal hit rates for workloads with skewed popularity distributions. Adapter switching takes approximately 50-200 microseconds per adapter on H100, dominated by PCIe transfer of the adapter weights.
S-LoRA (Stanford's system for serving thousands of LoRA adapters) improves on this through unified memory pooling and ahead-of-time preloading. Instead of loading individual adapters on demand, S-LoRA maintains a unified memory pool where adapter weights are staged on the GPU and swapped under a predict-then-preload policy. The system uses a popularity-based prefetcher that observes request patterns and preloads likely-to-be-requested adapters before they are needed. At 10,000 adapters with 200 concurrent requests, S-LoRA achieves 94% adapter cache hit rates compared to 82% for on-demand loading, reducing average per-request adapter loading latency from 120 microseconds to 15 microseconds.
| Architecture | Adapter Cache Hit Rate | Switching Latency |
|---|---|---|
| On-Demand Loading | 75-82% | 50-200 us |
| vLLM S-LoRA | 88-94% | 10-30 us |
| Predict-Then-Preload | 92-96% | 5-15 us |
| GPU-CPU Unified Memory | 95-98% | 2-5 us |
GPU MEMORY POOLING AND ADAPTER MANAGEMENT
The GPU memory budget for adapter weights is the key configuration parameter in multi-LoRA serving. For a 70B model consuming approximately 140 GB of GPU memory at FP16 across 8 H100s, roughly 10-20 GB of headroom remains for adapters on each GPU. With rank-16 adapters at 4 MB each, this allows 2,500-5,000 adapters in the on-GPU pool. Storing adapters in FP8 halves the memory requirement but degrades generation quality by 0.5-1.5% on domain-specific tasks. The remaining adapters reside in host DRAM, where 10,000 adapters consume only 40 GB of host memory.
Adapter weight quantization offers another optimization path. LoRA adapters can be quantized to INT4 with minimal quality loss for most fine-tuning tasks. At INT4, a rank-16 70B adapter shrinks to 2 MB, doubling the on-GPU adapter capacity. The quantization overhead during inference adds approximately 0.1 ms per LoRA layer due to dequantization, which is acceptable for latency-sensitive applications targeting sub-100ms per-token generation. For batch sizes above 8, the dequantization cost is amortized across the batch and becomes negligible.
CONCURRENT ADAPTER SERVING AND BATCHING
Multi-LoRA serving's core challenge is batching across requests with different adapters. A naive implementation runs separate batch iterations per adapter, fragmenting GPU utilization. vLLM's approach groups requests by adapter into sub-batches, running a single forward pass for all requests sharing the same adapter before switching to the next. This adapter-aware batching achieves 85-95% of single-adapter throughput when traffic is concentrated on the top 10-20 adapters, but degrades to 60-70% when traffic is evenly distributed across 100+ adapters.
Punica (a multi-LoRA serving system published at SOSP 2023, now integrated into vLLM) introduced a kernel design that processes multiple LoRA adapters simultaneously within a single CUDA kernel launch. Instead of serializing adapter execution, Punica's segmented matrix multiplication dispatches blocks from multiple adapters onto different SMs in a single kernel, achieving near-perfect utilization even with heterogeneous adapter ranks. Production deployments using Punica kernels achieve throughput of 3,800 tok/s for 256 concurrent adapters on a single H100 serving Llama 70B, compared to 1,200 tok/s with adapter-serial batching.
PRODUCTION DEPLOYMENT PATTERNS
The most common production pattern is popularity-tiered adapter placement. Identify the top 10% of adapters by request volume and pin them resident in GPU memory. The middle 30% cycle through a reserved memory buffer. The bottom 60% are loaded on demand from host DRAM with a secondary buffer in CPU memory. This three-tier pattern achieves 92-97% effective adapter cache hit rates while minimizing GPU memory waste for long-tail adapters that receive one request per day.
Adapter routing adds another dimension. For deployments across multiple GPU nodes, route requests for the same adapter to the same node through affinity-aware load balancing. This improves adapter cache locality and reduces inter-node adapter transfer. When a node fails, its adapter set must be redistributed - implement graceful degradation by replicating the top 20% of adapters across at least 2 nodes. For cost optimization on ClusterBid, the sweet spot for multi-LoRA serving is an 8x H100 node serving 500-2,000 adapters at $1.15/hr, achieving per-adapter inference cost of $0.08-0.12 per million tokens, which is 40-60% cheaper than deploying separate models per fine-tune.
| Configuration | Adapters Served | Cost per 1M Tokens |
|---|---|---|
| 1x H100 pinned + on-demand | 100-500 | $0.08-0.12 |
| 8x H100 pinned + on-demand | 500-2,000 | $0.04-0.07 |
| 8x H100 (S-LoRA unified) | 2,000-10,000 | $0.03-0.05 |
| Separate models per adapter | 10-50 | $0.35-0.89 |
| Savings vs separate models | 10-100x adapters | 40-80% cheaper |
