The Three SFT Approaches for DeepSeek V3: Full, LoRA, and QLoRA
Fine-tuning a 671B MoE model is not the same problem as fine-tuning a 70B dense model. The active-parameter count is 37B per token, but the memory requirements demand the full 671B footprint. Every training approach - Full SFT, LoRA, and QLoRA - makes a different tradeoff between GPU count, training time, and model quality. The gap between minimal and maximal config is wider than for any dense model on the market, and choosing wrong can mean overpaying by 20x for the same fine-tuning run.
Full SFT updates all 671B parameters. This requires the entire model in high precision (BF16) plus gradients and AdamW optimizer states - roughly 6-8 TB of VRAM in practice. That translates to 48-60 H200 SXM GPUs, or 64-80 H100 SXM GPUs, with tightly coupled InfiniBand for gradient synchronization across the MoE all-to-all fabric. LoRA freezes the base model and inserts low-rank adapters (rank 16-64) into attention and feed-forward projections. The base model still sits in VRAM at 1,342 GB for BF16, so you need 8-16 H200s, but the adapter parameters and optimizer states are negligible by comparison. QLoRA pushes the base model down to NF4 (4-bit NormalFloat), compressing 671B parameters to roughly 336 GB, and keeps LoRA adapters in BF16. This is the only approach that fits on 4-8 H200s.
The quality gap between these approaches is real but narrower than the cost gap. For domain adaptation (code, legal, medical), properly tuned LoRA at rank 64 matches Full SFT performance within 2-5% on domain-specific benchmarks. QLoRA adds 0.5-2% additional degradation over LoRA depending on dataset size and task complexity. The diminishing returns of Full SFT on a 671B model are steep when the training dataset is under 100K examples. For most production fine-tuning runs, LoRA on 8-16 GPUs is the correct default. For a broader comparison of MoE hardware requirements, see our MoE GPU infrastructure guide.
Full SFT: The 48-GPU Floor and What You Actually Pay Per Run
Training DeepSeek V3 from a checkpoint requires more than just fitting the weights. At BF16 mixed precision, the memory breakdown is: 1,342 GB for weights, 1,342 GB for gradients (computed in FP16 and accumulated in FP32), and 2,684 GB for AdamW states (momentum and variance at FP32). Add activation memory for a batch of 8 sequences at 2,048 tokens: roughly 150-200 GB for MoE activation storage given the expert routing overhead. Total: approximately 5.6 TB minimum, rounded up to 6+ TB with distributed training overhead, communication buffers, and ZeRO stage-2 fragments.
Forty-eight H200 SXM GPUs provide 6,768 GB total at $2.02/GPU/hr on-demand - that is $96.96/hr for the cluster. A full SFT run on 10,000 examples at 2,048 sequence length, with batch size 64 (8 per GPU with gradient accumulation), runs approximately 70-80 hours including validation checkpoints and warmup. Total cost: $6,800-$7,800 at on-demand rates. Scaling to 100K examples pushes this to 700+ hours and $68K-$78K. These numbers assume optimal MoE load balancing - expert dropout and imbalanced routing can add 20-40% to training time depending on how well your data distribution matches the pre-training distribution.
Full SFT on H100s is cheaper per GPU but requires more GPUs. Sixty-four H100 SXM GPUs ($1.15/hr each, $73.60/hr total) can match the 48 H200 config, but with roughly 15-20% longer wall-clock time due to H200's 4.8 TB/s memory bandwidth advantage over H100's 3.35 TB/s. The effective cost-per-run ends up similar - roughly $7,500-$8,500 for the same 10K example run. The real constraint is cluster availability: 48 contiguous H200s with validated InfiniBand for MoE all-to-all is harder to source than 64 H100s on short notice.
| Component | GPU Count | GPU Hours (10K ex) | Total Cost (On-Demand) |
|---|---|---|---|
| Full SFT (BF16, 48 H200) | 48 | 3,500-3,800 | $7,000-$7,800 |
| Full SFT (BF16, 64 H100) | 64 | 4,400-4,800 | $5,100-$5,500 |
| Full SFT (FP8 mixed, 32 H200) | 32 | 2,800-3,200 | $5,700-$6,500 |
LoRA Fine-Tuning: The Practical Production Path at 8x H200
LoRA on DeepSeek V3 keeps the base model frozen at BF16 and trains low-rank adapters injected into specific modules. The dominant cost is hosting the frozen base model: 1,342 GB for weights, plus some overhead for activations and forward-pass intermediates. Eight H200 SXM GPUs (1,128 GB total) cannot fit the full BF16 base model and any meaningful batch size. The practical minimum is 8 H200s running at FP8 mixed precision (671 GB for weights), which leaves approximately 450 GB for activations, KV cache, and the LoRA adapter gradients.
At rank 64 on attention QKV projections plus intermediate feed-forward layers (roughly 1.2B trainable parameters, or 0.18% of total), the LoRA memory overhead is approximately 2.4 GB for adapter weights and 9.6 GB for AdamW optimizer states. The training bottleneck shifts from memory to compute: LoRA forward-backward passes are slower per sample than inference because you must materialize and differentiate through adapter weights while keeping the frozen base graph warm. Expect roughly 1.6x the per-sample latency of inference throughput on the same hardware.
For 10,000 examples at 2,048 sequence length on 8 H200 SXM GPUs at FP8: run time is approximately 22-28 hours depending on target loss, learning rate schedule, and validation frequency. At $16.16/hr (8 x $2.02), total cost is $355-$450 for the full fine-tuning run. For 100K examples, expect 220-280 hours at $3,555-$4,520. This is the cost regime where LoRA becomes the obvious default for teams running periodic domain adaptation - below $5K per major fine-tuning cycle even at larger dataset sizes.
| Config | GPU Hours (10K ex) | Total Cost | Quality vs Full SFT |
|---|---|---|---|
| LoRA rank 16, 8 H200 FP8 | 18-22 hrs | $290-$355 | 90-93% match |
| LoRA rank 32, 8 H200 FP8 | 20-25 hrs | $323-$404 | 93-96% match |
| LoRA rank 64, 8 H200 FP8 | 22-28 hrs | $355-$452 | 95-98% match |
QLoRA: Extreme Memory Compression for Single-Node Training
QLoRA quantizes the frozen base model to NF4 (4-bit NormalFloat) while keeping LoRA adapters in BF16. This brings the DeepSeek V3 base from 1,342 GB to roughly 336 GB, opening the possibility of 4-8 GPU configurations that were previously impossible. The NF4 format is not a simple INT4 truncation - it uses a normalized float distribution that preserves dynamic range by fitting the weight distribution to a standard normal, then quantizing the tails asymmetrically. This matters for MoE models because gating weights have different statistical properties than expert weights, and NF4 handles this better than naive INT4.
Four H200 SXM GPUs (564 GB total) running QLoRA with NF4 base (336 GB) leaves 228 GB for KV cache, activations, LoRA adapter states, and CPU-GPU transfer buffers. At rank 64 on attention and feed-forward layers, this fits batch sizes of 4-8 per GPU at 2K context before hitting memory pressure. Training throughput on 4 H200s is roughly 30-35% slower per-sample than the 8 H200 LoRA config due to the dequantization overhead of NF4 weights on every forward pass. Expect 30-38 hours for 10K examples on 4 H200s versus 22-28 hours on 8 H200s.
The quality delta: QLoRA at rank 64 on DeepSeek V3 shows 1-3% degradation versus full-precision LoRA on domain-specific benchmarks (code, math, instruction following). The gap is largest on multi-step reasoning tasks where the NF4 quantization of gating weights compounds routing noise across steps. For classification, entity extraction, summarization, and RAG fine-tuning, the gap is under 1% and often within the variance of the training run itself. QLoRA on 4 H200s at $8.08/hr ($323-$430 for 10K examples) is the cheapest viable entry point for applying DeepSeek V3 to a specific domain.
| Approach | GPUs | GPU Hours (10K ex) | Total Cost | Quality vs Full SFT |
|---|---|---|---|---|
| QLoRA rank 64, 4 H200 NF4 | 4 | 120-152 | $323-$430 | 93-96% |
| LoRA rank 64, 8 H200 FP8 | 8 | 176-224 | $355-$452 | 95-98% |
| Full SFT, 48 H200 BF16 | 48 | 3,500-3,800 | $7,000-$7,800 | Baseline |
Dataset Size Scaling: GPU Hours from 1K to 100K Examples
Fine-tuning cost scales roughly linearly with dataset size for all three approaches, but the slope is dramatically different. Full SFT requires expensive large-cluster hours for every sample processed. LoRA and QLoRA shift the cost to the forward pass through the frozen base model, which is 8-16 GPUs regardless of whether you process 1K or 100K examples. The table below shows realistic run times and total costs for a LoRA rank 64 config on 8 H200 SXM GPUs ($2.02/GPU/hr) at 2,048 sequence length, assuming optimal data loading and no I/O bottlenecks.
The per-example cost for LoRA drops significantly at larger dataset sizes because the one-time overhead of model loading, compilation, validation checkpoints, and warmup is amortized. A 1K-example run costs roughly $80-$100 with a 5-7 hour wall time. A 100K-example run costs $3,500-$4,500 with 220-280 hours. The marginal cost per additional 1K examples settles at roughly $30-$35 after the first 10K, making LoRA increasingly efficient as your dataset grows relative to the fixed overhead of cluster setup and teardown.
The dataset quality inflection point for DeepSeek V3 LoRA is around 3,000-5,000 high-quality examples. Below this threshold, the adapter tends to memorize rather than generalize, and you see minimal improvement on held-out evaluation datasets. Above 10K examples, improvements continue but with diminishing returns. The practical recommendation: start with 3K-5K curated examples on LoRA rank 32 at $140-$200 to validate domain fit, then scale to 10K-20K at rank 64 for production deployment. Full SFT only makes financial sense above 50K-100K examples where LoRA adapter capacity (rank 64) starts to saturate.
| Examples | Full SFT Hours (48 H200) | LoRA Hours (8 H200) | LoRA Total Cost |
|---|---|---|---|
| 1,000 | 400-450 | 5-7 | $80-$113 |
| 5,000 | 1,800-2,000 | 12-16 | $194-$258 |
| 10,000 | 3,500-3,800 | 22-28 | $355-$452 |
| 50,000 | 17,500-19,000 | 110-140 | $1,780-$2,260 |
| 100,000 | 35,000-38,000 | 220-280 | $3,555-$4,520 |
H200 vs H100: The Training Cost Showdown for DeepSeek V3 SFT
H200 SXM (141 GB VRAM, 4.8 TB/s bandwidth) versus H100 SXM (80 GB VRAM, 3.35 TB/s bandwidth) - the choice matters differently for training than for inference. During training, memory bandwidth dominates per-sample throughput because every forward and backward pass reads the full weight matrix. The H200's 1.45 TB/s additional bandwidth translates to roughly 30-35% faster training throughput for memory-bound operations. For compute-bound operations (large batch matmuls on tensor cores), the gap narrows to 10-15%. LoRA training on DeepSeek V3 is memory-bandwidth-bound because the frozen base weights dominate the memory traffic pattern.
On a per-dollar basis at on-demand rates: 8 H200s at $16.16/hr deliver ~35% more throughput than 8 H100s at $9.20/hr, but cost 76% more. The pure cost-efficiency winner for LoRA training is H100 at $1.15/GPU/hr if your batch size and sequence length fit within 80 GB per GPU. For DeepSeek V3 LoRA at FP8, the base model is 671 GB, which requires at minimum 8 H100s (640 GB) at FP8 with gradient checkpointing to fit, versus 8 H200s (1,128 GB) at FP8 which provides 457 GB of headroom. The H100 config at 640 GB is trivially too tight for any batch size above 4 at 2K context. The practical minimum for H100 LoRA is 12 H100s (960 GB) at $13.80/hr.
B200 enters the picture at $3.36/GPU/hr with 192 GB VRAM and FP4 tensor core support. For LoRA training, the additional VRAM (192 GB vs 141 GB) is not the primary constraint - DeepSeek V3 already fits on 8 H200s. The FP4 tensor cores cannot accelerate frozen-base forward passes at FP8, so B200 does not deliver a meaningful training throughput advantage over H200 for LoRA workloads. The exception is Full SFT at FP4 mixed precision, where B200's memory capacity (192 GB per GPU) and FP4 throughput combine to reduce the Full SFT GPU count from 48 H200s to roughly 32 B200s, at $107.52/hr versus $96.96/hr - slightly more expensive per hour but with shorter wall time. Pricing reflects June 2026 on-demand market rates and can fluctuate with GPU availability.
| GPU | VRAM | Bandwidth | On-Demand Rate | Min LoRA Config |
|---|---|---|---|---|
| H100 SXM5 | 80 GB | 3.35 TB/s | $1.15/hr | 12x ($13.80/hr) |
| H200 SXM | 141 GB | 4.80 TB/s | $2.02/hr | 8x ($16.16/hr) |
| B200 SXM6 | 192 GB | 8.00 TB/s | $3.36/hr | 6x ($20.16/hr) |
Rent vs Reserve: Procurement Strategy for Fine-Tuning Runs
Fine-tuning workloads have a different procurement profile than inference. Inference runs 24/7 and demands low latency - reserved contracts with guaranteed capacity make sense. Fine-tuning runs are finite (days to weeks), interruptible with checkpointing, and rarely require the exact same cluster size from one run to the next. On-demand rental is the correct default for all three SFT approaches unless you are running continuous training pipelines with weekly checkpoint cycles.
Spot instances are viable for LoRA and QLoRA fine-tuning if your framework supports checkpointing and resumption. DeepSeek V3 training with DeepSpeed ZeRO stage 2 or 3 writes optimizer state and adapter checkpoints in under two minutes on 8-16 GPU clusters. Spot eviction on H100 and H200 clusters in mid-2026 averages 1-3 interruptions per week on the cheapest spot tier, and the cost savings versus on-demand is typically 50-65%. The effective net cost for a LoRA run on 8 H200s drops from $16.16/hr to roughly $5.65-$8.08/hr on a well-configured spot strategy with 30-minute checkpoint intervals.
When reserved contracts make sense: three specific scenarios. First, continuous SFT pipelines where you fine-tune weekly or biweekly on new data - a 3-month reserved rate on 8 H200s typically locks in 25-35% below on-demand. Second, research teams running iterative experiments (learning rate sweeps, rank searches, dataset ablations) that keep the cluster warm between runs. Third, any team doing Full SFT runs above 50K examples where the total on-demand cost exceeds $40K per run and the schedule is known in advance. For all other fine-tuning scenarios, on-demand or spot through a marketplace like ClusterBid provides the flexibility to resize between Full SFT, LoRA, and QLoRA configs without being locked into GPU counts that don't match the specific run. Browse ClusterBid's live inventory for current H100, H200, and B200 availability and pricing.
| Procurement Model | Effective GPU/hr (H200) | Best For |
|---|---|---|
| On-Demand | $2.02 | Single runs, experiments, <10K ex |
| Spot (with checkpointing) | $0.71-$1.01 | LoRA runs, interruptible SFT |
| 3-Month Reserved | $1.31-$1.52 | Continuous training pipelines |
| 12-Month Reserved | $1.01-$1.21 | Full SFT, sustained clusters |
