All essays
BenchmarkCOMPARISONFEB 2026

Adapter-Based Fine-Tuning: LoRA, DoRA, AdaLoRA GPU Memory vs Accuracy Tradeoffs

Deep-dive comparison of adapter-based fine-tuning methods: LoRA, DoRA, AdaLoRA, and PiSSA. GPU memory benchmarks, accuracy vs parameter count tradeoffs, multi-adapter serving infrastructure, and deployment strategies for production fine-tuning at scale.

01

THE PARAMETER-EFFICIENT FINE-TUNING LANDSCAPE

Adapter-based fine-tuning methods add a small number of trainable parameters while freezing the base model, reducing optimizer state memory by 95-99 percent. LoRA (Low-Rank Adaptation) decomposes weight updates as W' = W + BA, where B and A are low-rank matrices. DoRA applies the same decomposition to the directional component while learning a separate magnitude vector. AdaLoRA dynamically allocates rank across layers based on importance scores. PiSSA initializes adapters with the principal singular components of the base weights rather than random initialization.

The core tradeoff is parameter count versus accuracy. A rank-16 LoRA on all query and value projection matrices of a 70B model adds 84M trainable parameters (0.12 percent of total). DoRA doubles this to 168M because it adds a learnable magnitude vector per layer. AdaLoRA with an average rank of 16 uses similar parameter counts but concentrates rank on attention layers where it matters most. All three deliver 90-99 percent of full fine-tuning accuracy at 0.1-2 percent of the trainable parameter count.

MethodTrainable Params (70B)Training VRAM (70B)Speed vs Full FTAccuracy vs Full FTBest For
Full Fine-Tune70B (100%)560-700 GB1x (baseline)100%Maximum quality
LoRA r=842M (0.06%)140-160 GB2.5-3x faster96-98%Most efficient
LoRA r=64336M (0.48%)145-170 GB2-2.5x faster98-99.5%Quality balanced
DoRA r=16168M (0.24%)142-165 GB2-2.5x faster98-99%Direction learning
AdaLoRA80-160M (0.15%)145-170 GB1.8-2x faster98-99.5%Adaptive rank
PiSSA r=1684M (0.12%)140-160 GB2.5-3x faster98-99%Better init
02

GPU MEMORY BREAKDOWN: WHERE THE BYTES GO

For LoRA training of a 70B model on a single H100 with ZeRO-3, memory breaks down as follows: base model weights distributed across GPUs (140 GB / 8 GPUs = 17.5 GB per GPU with ZeRO-3), LoRA weights (168 MB total, negligible), optimizer states for LoRA weights (84M * 12 bytes = 1 GB for Adam states and gradients), activations (batch size 1, 8K sequence = 8-12 GB with activation checkpointing), and KV cache for generation (2-4 GB). Total per-GPU: approximately 30-35 GB, fitting comfortably on a single H100.

The surprising finding: activation memory dominates, not adapter weights. LoRA reduces optimizer memory by >99 percent, but activation memory remains identical to full fine-tuning because activations depend on the sequence length and batch size, not the parameter count. For 8K context with batch size 4, activation memory reaches 40-45 GB per GPU even with checkpointing. This is why AdaLoRA's adaptive rank allocation that reduces activation recomputation at unimportant layers can save 15-25 percent activation memory - a benefit independent of parameter count.

Memory ComponentFull FT 70B (8 H100)LoRA 70B (8 H100)LoRA 70B (1 H100)QLoRA 70B (1 H100)
Base Weights140 GB (17.5/GPU)140 GB (17.5/GPU)140 GB17.5 GB (NF4)
Optimizer States280 GB (35/GPU)1 GB (0.12/GPU)1 GB1 GB
Gradients140 GB (17.5/GPU)168 MB168 MB168 MB
Activations (BS=1, 8K)8-12 GB/GPU8-12 GB/GPU8-12 GB8-12 GB
Total per GPU~70 GB/GPU~30 GB/GPU~150 GB (>80 GB!)~28 GB
Max Batch Size (8K ctx)4-88-161 (OOM at 2)4-8
03

SERVING FINE-TUNED ADAPTERS IN PRODUCTION

Adapter serving in production uses a base model loaded once in GPU memory and switches adapter weights per request. The base model weights (140 GB for 70B at FP16) are static. Each LoRA adapter (rank 16) adds 168 MB for parameter loading plus 200-500 MB for KV cache overhead from the merged output. With 200 adapters on a single H100, total memory is 140 GB (base) + 35 GB (adapters) + 12 GB (KV cache) = 187 GB, which exceeds 80 GB. QLoRA base at NF4 reduces to 17.5 GB, allowing 200 adapters + KV cache in 80 GB.

The critical deployment metric is adapter switching throughput. With dynamic adapter loading from CPU RAM, switching takes 2-5 ms with the base model staying resident. vLLM, TensorRT-LLM, and SGLang all support LoRA serving with different tradeoffs: vLLM's Punica kernel achieves zero-overhead switching for up to 256 adapters. Production FTaaS platforms like Predibase serve 5,000+ adapters per 8-GPU node using this architecture.

MetricLoRA r=16LoRA r=64DoRA r=16AdaLoRA
Adapter Size (70B)168 MB672 MB336 MB160-320 MB
Cold Switch2-5 ms5-12 ms3-6 ms2-5 ms
Merge Time (H100)15-30s60-120s20-35s15-30s
Adapters/GPU (FP16)0000
Adapters/GPU (NF4)~200~50~100~200
Inference Overhead2-5%8-15%3-6%2-5%
Quality (MMLU 70B)82.3%83.1%82.8%83.2%
04

ADALORA: WHY ADAPTIVE RANK MATTERS FOR GPU MEMORY

AdaLoRA parameterizes weight updates as P * diag(s) * Q, where s is a learnable vector of singular values. During training, unimportant singular values are pruned based on their contribution to the loss, effectively allocating higher rank to attention layers and lower rank to MLP layers. In practice, AdaLoRA allocates 70-80 percent of total rank budget to attention projections and only 20-30 percent to feed-forward layers, matching the empirical finding that attention fine-tuning matters more.

The GPU memory impact is twofold. First, AdaLoRA's adaptive pruning means the effective rank during inference is 60-70 percent of the nominal rank, reducing adapter storage by 30-40 percent. Second, the SVD parameterization is more memory-efficient during training because the diagonal s matrix requires only O(k) parameters versus O(2 * d * k) for LoRA's BA decomposition at rank k. For a 70B model, AdaLoRA with average rank 16 uses 110-140 MB per adapter versus 168 MB for LoRA, a 17-35 percent reduction that accumulates significantly across thousands of adapters.

05

MULTI-GPU ADAPTER TRAINING: SCALING FINE-TUNING INFRASTRUCTURE

Adapter fine-tuning across multiple GPUs uses the same parallelism strategies as full fine-tuning but with drastically reduced communication volume. With ZeRO-3 on 8 GPUs and LoRA, only the LoRA weights (84M parameters) and their optimizer states are sharded - roughly 0.1 percent of the communication of full fine-tuning. This means LoRA training achieves near-linear scaling with GPU count up to 64 GPUs, versus the 32-GPU saturation point for full fine-tuning.

The practical deployment: a 16-GPU H100 cluster fine-tunes 200 adapters concurrently using a round-robin scheduler. Each adapter trains on 1-4 GPUs for 30-120 minutes. The training infrastructure must handle GPU allocation, adapter checkpointing (168 MB each), and merge queue management. Total training throughput: 50-100 adapters per day on a 16-GPU cluster. At $3.00/GPU-hour, per-adapter training cost is $3-12.

06

B200 AND THE ADAPTER ECONOMICS REVOLUTION

B200's 192 GB VRAM enables QLoRA training of 70B models with batch size 8-16 at 8K context - matching H100's batch size with 2-3x the throughput per GPU. Training a rank-16 LoRA on a 70B model on B200 completes in 20-30 minutes versus 45-60 minutes on H100. For a service fine-tuning 1,000 adapters per month, B200 reduces training GPU-hours from 2,000 to 800, saving $3,600/month.

In serving configuration, a single B200 hosts 500+ NF4 adapters (versus 200 on H100) and switches between them with negligible overhead. For a FTaaS deployment with 10,000 adapters, B200 reduces the serving GPU count from 50 to 20, cutting serving infrastructure cost by 55-60 percent at equivalent throughput.

Filed under
LoRA Fine TuningDoRA Adapter MethodAdaLoRA GPU MemoryAdapter Training GPUPEFT ComparisonLow Rank AdaptationGPU Fine Tuning