THE FINE-TUNING-AS-A-SERVICE LANDSCAPE
FTaaS platforms (Together AI, Fireworks AI, Anyscale) let a single base model be customized for thousands of customers through PEFT techniques, with each adapter consuming only 0.1-2 percent of the base model's parameter count. The challenge is serving 10,000+ adapters on shared GPUs with sub-second cold-start latency.
At $2.50-4.00 per GPU-hour and $10-20 per adapter per month revenue, a deployment needs 50-100 adapters per GPU to achieve positive margins. This drives specialized adapter routing systems that minimize memory cost per adapter and maximize GPU utilization.
LORA AND QLORA: THE ADAPTER INFRASTRUCTURE STACK
A rank-16 LoRA adapter for Llama-3.1-70B adds 84M parameters (168 MB at FP16), 0.24 percent of the base model's 140 GB. Base model weights are stored once in GPU memory; adapter weights load on demand per request. QLoRA extends this with 4-bit NF4 quantization of the base model, reducing memory from 140 GB to 17.5 GB with 1-3 percent quality degradation.
A single H100 with QLoRA holds: base model (17.5 GB), KV cache for 32 requests (12 GB), and approximately 200 adapters (33.6 GB), totaling 63.1 GB. At 4 GPUs per node, one node serves 800 adapters with 32 concurrent users. Scaling to 10,000 adapters requires 13 nodes (52 H100 GPUs).
| Config | LoRA FP16 Base | QLoRA NF4 | LoRA FP8 Base | DoRA FP16 |
|---|---|---|---|---|
| Base Mem (70B) | 140 GB | 17.5 GB | 70 GB | 140 GB |
| Adapter Mem r16 | 168 MB | 168 MB | 168 MB | 336 MB |
| Max Adapter/H100 | 0 | ~200 | ~50 | 0 |
| MMLU Quality | 85.0% | 82.5% | 84.2% | 85.3% |
| Train VRAM | 140 GB | 24 GB | 48 GB | 140 GB |
| Adapter Switch | 2-5 ms | 2-5 ms | 2-5 ms | 2-5 ms |
| Best For | Quality | Density | Balance | Training only |
ADAPTER ROUTING: THE INFERENCE GATEWAY
The adapter routing layer maps each request to its adapter weights before inference. Three approaches: static allocation (adapters loaded at startup, simplest but limited count), dynamic loading (from CPU on demand, supports unlimited adapters but adds 2-5 ms transfer), and hybrid tiered (hot set in GPU memory, cold loaded from CPU).
Production FTaaS platforms use the hybrid approach: a hot set of 500-1000 adapters in GPU memory distributed via consistent hashing. With LRU eviction and 10,000 adapters, hot cache achieves 70-85 percent hit rate at 1000 slots and >95 percent at 2000 slots. The routing gateway uses Redis or etcd for adapter-to-GPU location metadata.
| Strategy | Max Adapters | Hot Latency | Cold Latency | Hit Rate 10K | GPU Overhead |
|---|---|---|---|---|---|
| Static Alloc | 200/GPU | 2-5 us | N/A | 100% | High |
| Dynamic Load | Unlimited | 2-5 us | 2-5 ms | Ad hoc | Zero |
| Hybrid 1K slot | Unlimited | 3-10 us | 2-5 ms | 70-85% | 168 MB/slot |
| Hybrid 2K slot | Unlimited | 3-10 us | 2-5 ms | >95% | 336 MB/slot |
MULTI-TENANT FINE-TUNING: TRAINING INFRASTRUCTURE
Training GPUs are separate from serving GPUs due to different compute profiles requiring full backward passes and optimizer states. For QLoRA training of Llama-3.1-70B, a single H100 trains with batch size 4, gradient accumulation 8, and sequence length 8192. Training takes 30-120 minutes on 1-4 GPUs per adapter.
The scheduling challenge is packing variable-length training jobs. A bin-packing scheduler groups jobs by GPU requirement and duration. Training GPU cost per adapter: 1 GPU-hour at $2.50-3.50. At $10-20 per adapter per month revenue, payback is 1-2 weeks. Serving cost adds $0.0003-0.001 per query, negligible compared to training cost.
ADAPTER MERGING AND WEIGHT COMPOSITION
Adapter merging combines LoRA weights into the base model: W' = W + s * BA for all linear layers. Merging eliminates adapter loading at inference, enabling standard vLLM/TGI serving with zero adapter overhead. A merge for rank-16 into Llama-70B requires 2.5 TFLOPS and takes 15-30 seconds on an H100.
Multi-adapter composition (domain + style + safety adapters) using TIES-Merging or DARE algorithms enables a marketplace of composable adapters. The infrastructure is a merge service that loads the base model once, applies N adapters sequentially, and produces a static merged model.
Batch merging 100 adapters requires a dedicated GPU pool processing 4-8 merges simultaneously, with validation before deployment. Total pipeline: 45-90 seconds per merged adapter.
B200 AND THE ECONOMICS OF FTaaS
A single B200 in QLoRA mode holds 600 adapters versus H100's 200. This 3x density directly reduces per-adapter serving cost by 3x. For 10,000 adapters, GPU count drops from 52 H100 to 17 B200, reducing infrastructure cost by 55-60 percent.
B200 also accelerates training: a QLoRA adapter on Llama-70B completes in 35-45 minutes versus 60-90 minutes on H100 (40-50 percent reduction). Compound effect: B200-based FTaaS achieves approximately 2.5-3x adapter throughput per dollar compared to H100.
