All essays
TechnicalDEEP DIVEFEB 2026

Fine-Tuning Infrastructure: GPU Cluster Design for Model Customization

Infrastructure design for LLM fine-tuning covering parameter-efficient methods (LoRA, QLoRA), full fine-tuning cluster topology, data pipeline, and hyperparameter optimization at scale.

01

FINE-TUNING APPROACHES AND GPU REQUIREMENTS

Fine-tuning spans three approaches with vastly different GPU requirements. Full fine-tuning updates all model parameters requiring 4x model size in GPU memory (480 GB for Llama 3 70B with AdamW). LoRA fine-tuning trains 0.1-1 percent of parameters using rank-64 adapters, requiring 1.2x model size (170 GB for Llama 3 70B). QLoRA combines LoRA with INT4 quantization, fitting Llama 3 70B on a single H100 with 80 GB.

Training throughput comparison: full fine-tuning Llama 3 8B on 8 H100 GPUs achieves 4,500 tokens/second with batch size 64. LoRA on the same hardware achieves 5,800 tokens/second (29 percent faster). QLoRA with INT4 base model achieves 3,200 tokens/second but enables 70B fine-tuning on 1 GPU versus 8 GPUs for full fine-tuning. Cost per fine-tuning run: LoRA $200 (8 GPU-hours) vs full $1,600 (64 GPU-hours) for Llama 3 8B.

MethodGPU Memory (70B)Min GPUsThroughput (70B)Quality vs Full FTCost per Run
Full fine-tuning~480 GB8 H100100 tok/s/GPUBaseline$12,000-$18,000
LoRA (rank=64)~170 GB2 H100180 tok/s/GPU95-98%$2,000-$4,000
QLoRA (INT4 + LoRA)~48 GB1 H10085 tok/s/GPU93-96%$800-$1,500
DoRA (LoRA + adapters)~175 GB2 H100170 tok/s/GPU96-99%$2,500-$4,500
Full + DeepSpeed ZeRO-3~250 GB4 H100140 tok/s/GPU100%$6,000-$10,000
02

DATA PIPELINE FOR FINE-TUNING

Fine-tuning data quality determines model quality more than method choice. A data curation pipeline: raw data ingestion (100 GB-10 TB), deduplication using MinHash (removes 15-30 percent duplicates), quality filtering via perplexity-based scoring (filters bottom 10 percent), prompt-response pair extraction, and format standardization. Pipeline throughput must exceed training consumption rate: 1,000-10,000 tokens/second per GPU.

Data loading optimization critical for fine-tuning throughput. WebDataset format with sharded tar files achieves 90+ percent storage bandwidth utilization versus 40-60 percent for individual file reads. NVIDIA DALI preprocessing pipeline offloads data augmentation to GPU, reducing CPU bottleneck. With DALI, data loading completes in 2-5ms per batch versus 15-40ms without, eliminating pipeline stalls.

03

HYPERPARAMETER OPTIMIZATION AT SCALE

Hyperparameter optimization for fine-tuning searches learning rate (1e-5 to 5e-5), LoRA rank (8-256), LoRA alpha (8-256), warmup steps (0-500), and weight decay (0.0-0.1). A grid search with 50 combinations at 8 GPUs per run consumes 400 GPU-hours per model. Bayesian optimization reduces this to 15-25 runs (120-200 GPU-hours) achieving equivalent final performance.

Automated HPO infrastructure requires: parameter space definition, parallel trial execution across GPU partitions, early stopping for underperforming trials after 20 percent of training, and best run tracking. Optuna and Ray Tune integrate with SLURM/Kubernetes managing trial distribution. For a 128-GPU cluster, running 16 parallel HPO trials with 8 GPUs each optimizes resource utilization at 90-95 percent.

04

RLHF AND PREFERENCE TUNING INFRASTRUCTURE

RLHF training requires 4 model instances: policy model, reference model, reward model, and value model (for PPO). A 70B RLHF run requires 32-64 H100 GPUs versus 8 for SFT. Memory requirements reach 800 GB-1.2 TB due to simultaneous model loading. DeepSpeed Chat reduces memory via ZeRO-3 and LoRA-based PPO, fitting on 8-16 H100 GPUs.

Preference data collection infrastructure: 20-100 human annotators rating 1,000-10,000 prompt-response pairs daily. Annotation platform integrates with training pipeline, providing weekly reward model updates. DPO (Direct Preference Optimization) eliminates the need for separate reward and value models, reducing RLHF-style training GPU requirement by 40-50 percent while matching PPO quality.

Filed under
Fine-TuningLoRAQLoRAParameter-EfficientFull Fine-TuningSFTRLHFGPU Training