FINE-TUNING APPROACHES AND GPU REQUIREMENTS
Fine-tuning spans three approaches with vastly different GPU requirements. Full fine-tuning updates all model parameters requiring 4x model size in GPU memory (480 GB for Llama 3 70B with AdamW). LoRA fine-tuning trains 0.1-1 percent of parameters using rank-64 adapters, requiring 1.2x model size (170 GB for Llama 3 70B). QLoRA combines LoRA with INT4 quantization, fitting Llama 3 70B on a single H100 with 80 GB.
Training throughput comparison: full fine-tuning Llama 3 8B on 8 H100 GPUs achieves 4,500 tokens/second with batch size 64. LoRA on the same hardware achieves 5,800 tokens/second (29 percent faster). QLoRA with INT4 base model achieves 3,200 tokens/second but enables 70B fine-tuning on 1 GPU versus 8 GPUs for full fine-tuning. Cost per fine-tuning run: LoRA $200 (8 GPU-hours) vs full $1,600 (64 GPU-hours) for Llama 3 8B.
| Method | GPU Memory (70B) | Min GPUs | Throughput (70B) | Quality vs Full FT | Cost per Run |
|---|---|---|---|---|---|
| Full fine-tuning | ~480 GB | 8 H100 | 100 tok/s/GPU | Baseline | $12,000-$18,000 |
| LoRA (rank=64) | ~170 GB | 2 H100 | 180 tok/s/GPU | 95-98% | $2,000-$4,000 |
| QLoRA (INT4 + LoRA) | ~48 GB | 1 H100 | 85 tok/s/GPU | 93-96% | $800-$1,500 |
| DoRA (LoRA + adapters) | ~175 GB | 2 H100 | 170 tok/s/GPU | 96-99% | $2,500-$4,500 |
| Full + DeepSpeed ZeRO-3 | ~250 GB | 4 H100 | 140 tok/s/GPU | 100% | $6,000-$10,000 |
DATA PIPELINE FOR FINE-TUNING
Fine-tuning data quality determines model quality more than method choice. A data curation pipeline: raw data ingestion (100 GB-10 TB), deduplication using MinHash (removes 15-30 percent duplicates), quality filtering via perplexity-based scoring (filters bottom 10 percent), prompt-response pair extraction, and format standardization. Pipeline throughput must exceed training consumption rate: 1,000-10,000 tokens/second per GPU.
Data loading optimization critical for fine-tuning throughput. WebDataset format with sharded tar files achieves 90+ percent storage bandwidth utilization versus 40-60 percent for individual file reads. NVIDIA DALI preprocessing pipeline offloads data augmentation to GPU, reducing CPU bottleneck. With DALI, data loading completes in 2-5ms per batch versus 15-40ms without, eliminating pipeline stalls.
HYPERPARAMETER OPTIMIZATION AT SCALE
Hyperparameter optimization for fine-tuning searches learning rate (1e-5 to 5e-5), LoRA rank (8-256), LoRA alpha (8-256), warmup steps (0-500), and weight decay (0.0-0.1). A grid search with 50 combinations at 8 GPUs per run consumes 400 GPU-hours per model. Bayesian optimization reduces this to 15-25 runs (120-200 GPU-hours) achieving equivalent final performance.
Automated HPO infrastructure requires: parameter space definition, parallel trial execution across GPU partitions, early stopping for underperforming trials after 20 percent of training, and best run tracking. Optuna and Ray Tune integrate with SLURM/Kubernetes managing trial distribution. For a 128-GPU cluster, running 16 parallel HPO trials with 8 GPUs each optimizes resource utilization at 90-95 percent.
RLHF AND PREFERENCE TUNING INFRASTRUCTURE
RLHF training requires 4 model instances: policy model, reference model, reward model, and value model (for PPO). A 70B RLHF run requires 32-64 H100 GPUs versus 8 for SFT. Memory requirements reach 800 GB-1.2 TB due to simultaneous model loading. DeepSpeed Chat reduces memory via ZeRO-3 and LoRA-based PPO, fitting on 8-16 H100 GPUs.
Preference data collection infrastructure: 20-100 human annotators rating 1,000-10,000 prompt-response pairs daily. Annotation platform integrates with training pipeline, providing weekly reward model updates. DPO (Direct Preference Optimization) eliminates the need for separate reward and value models, reducing RLHF-style training GPU requirement by 40-50 percent while matching PPO quality.
