All essays
GuideGUIDEFEB 2026

Fine-Tuning as a Service: LoRA and QLoRA Deployment Patterns for Personalized Models

Production deployment patterns for fine-tuning services: LoRA, QLoRA, DoRA serving architectures. GPU requirements for multi-tenant fine-tuning, adapter routing, merge inference, and cost modeling for FTaaS platforms at scale.

01

THE FINE-TUNING-AS-A-SERVICE LANDSCAPE

FTaaS platforms (Together AI, Fireworks AI, Anyscale) let a single base model be customized for thousands of customers through PEFT techniques, with each adapter consuming only 0.1-2 percent of the base model's parameter count. The challenge is serving 10,000+ adapters on shared GPUs with sub-second cold-start latency.

At $2.50-4.00 per GPU-hour and $10-20 per adapter per month revenue, a deployment needs 50-100 adapters per GPU to achieve positive margins. This drives specialized adapter routing systems that minimize memory cost per adapter and maximize GPU utilization.

02

LORA AND QLORA: THE ADAPTER INFRASTRUCTURE STACK

A rank-16 LoRA adapter for Llama-3.1-70B adds 84M parameters (168 MB at FP16), 0.24 percent of the base model's 140 GB. Base model weights are stored once in GPU memory; adapter weights load on demand per request. QLoRA extends this with 4-bit NF4 quantization of the base model, reducing memory from 140 GB to 17.5 GB with 1-3 percent quality degradation.

A single H100 with QLoRA holds: base model (17.5 GB), KV cache for 32 requests (12 GB), and approximately 200 adapters (33.6 GB), totaling 63.1 GB. At 4 GPUs per node, one node serves 800 adapters with 32 concurrent users. Scaling to 10,000 adapters requires 13 nodes (52 H100 GPUs).

ConfigLoRA FP16 BaseQLoRA NF4LoRA FP8 BaseDoRA FP16
Base Mem (70B)140 GB17.5 GB70 GB140 GB
Adapter Mem r16168 MB168 MB168 MB336 MB
Max Adapter/H1000~200~500
MMLU Quality85.0%82.5%84.2%85.3%
Train VRAM140 GB24 GB48 GB140 GB
Adapter Switch2-5 ms2-5 ms2-5 ms2-5 ms
Best ForQualityDensityBalanceTraining only
03

ADAPTER ROUTING: THE INFERENCE GATEWAY

The adapter routing layer maps each request to its adapter weights before inference. Three approaches: static allocation (adapters loaded at startup, simplest but limited count), dynamic loading (from CPU on demand, supports unlimited adapters but adds 2-5 ms transfer), and hybrid tiered (hot set in GPU memory, cold loaded from CPU).

Production FTaaS platforms use the hybrid approach: a hot set of 500-1000 adapters in GPU memory distributed via consistent hashing. With LRU eviction and 10,000 adapters, hot cache achieves 70-85 percent hit rate at 1000 slots and >95 percent at 2000 slots. The routing gateway uses Redis or etcd for adapter-to-GPU location metadata.

StrategyMax AdaptersHot LatencyCold LatencyHit Rate 10KGPU Overhead
Static Alloc200/GPU2-5 usN/A100%High
Dynamic LoadUnlimited2-5 us2-5 msAd hocZero
Hybrid 1K slotUnlimited3-10 us2-5 ms70-85%168 MB/slot
Hybrid 2K slotUnlimited3-10 us2-5 ms>95%336 MB/slot
04

MULTI-TENANT FINE-TUNING: TRAINING INFRASTRUCTURE

Training GPUs are separate from serving GPUs due to different compute profiles requiring full backward passes and optimizer states. For QLoRA training of Llama-3.1-70B, a single H100 trains with batch size 4, gradient accumulation 8, and sequence length 8192. Training takes 30-120 minutes on 1-4 GPUs per adapter.

The scheduling challenge is packing variable-length training jobs. A bin-packing scheduler groups jobs by GPU requirement and duration. Training GPU cost per adapter: 1 GPU-hour at $2.50-3.50. At $10-20 per adapter per month revenue, payback is 1-2 weeks. Serving cost adds $0.0003-0.001 per query, negligible compared to training cost.

05

ADAPTER MERGING AND WEIGHT COMPOSITION

Adapter merging combines LoRA weights into the base model: W' = W + s * BA for all linear layers. Merging eliminates adapter loading at inference, enabling standard vLLM/TGI serving with zero adapter overhead. A merge for rank-16 into Llama-70B requires 2.5 TFLOPS and takes 15-30 seconds on an H100.

Multi-adapter composition (domain + style + safety adapters) using TIES-Merging or DARE algorithms enables a marketplace of composable adapters. The infrastructure is a merge service that loads the base model once, applies N adapters sequentially, and produces a static merged model.

Batch merging 100 adapters requires a dedicated GPU pool processing 4-8 merges simultaneously, with validation before deployment. Total pipeline: 45-90 seconds per merged adapter.

06

B200 AND THE ECONOMICS OF FTaaS

A single B200 in QLoRA mode holds 600 adapters versus H100's 200. This 3x density directly reduces per-adapter serving cost by 3x. For 10,000 adapters, GPU count drops from 52 H100 to 17 B200, reducing infrastructure cost by 55-60 percent.

B200 also accelerates training: a QLoRA adapter on Llama-70B completes in 35-45 minutes versus 60-90 minutes on H100 (40-50 percent reduction). Compound effect: B200-based FTaaS achieves approximately 2.5-3x adapter throughput per dollar compared to H100.

Filed under
Fine-Tuning as a ServiceLoRA DeploymentQLoRA ServingGPU Fine-TuningAdapter RoutingFTaaS InfrastructureMulti-Tenant Fine-Tuning