All essays
TechnicalDEEP DIVEFEB 2026

LLM Fine-Tuning Infrastructure: Dataset Management, Experiment Tracking, and GPU Orchestration

Full-stack infrastructure guide for LLM fine-tuning at scale. DVC and HuggingFace Datasets for versioning, WandB vs MLflow, GPU orchestration for LoRA/QLoRA jobs with 2026 cost examples.

01

Fine-Tuning at Scale Is an Infrastructure Problem First

A single LoRA fine-tuning run on Llama 3.1 70B with 10,000 instruction examples requires about 120 GPU-minutes on an H100 at roughly $5.70 in compute cost (at ClusterBid's June 2026 H100 rate of $2.85 per GPU-hour). Add hyperparameter search across 32 combinations (learning rate, rank, alpha, dropout, warmup steps) and you are at $182 and 32 hours of wall-clock time. Add dataset version iterations (new examples, corrected labels, domain splits) across 8 experiment cycles, and the total is $1,456 and roughly 26 GPU-days for a single model's fine-tuning campaign. The infrastructure cost of fine-tuning is not the individual training run. It is the combinatorial explosion of runs required by exploration.

The three layers of fine-tuning infrastructure that determine whether that $1,456 campaign produces useful results or is mostly wasted compute are dataset management (ensuring every run can be reproduced from a specific data snapshot), experiment tracking (recording hyperparameters, metrics, and environment metadata for every run), and GPU orchestration (scheduling many runs efficiently on shared GPU resources). Each layer independently affects the cost per useful experiment, and the three layers interact: a GPU orchestration system that cannot co-locate multiple LoRA runs on the same GPU wastes 50-70% of available throughput, while experiment tracking that does not capture the dataset version hash makes every successful run hard to reproduce and every failed run impossible to debug.

Most teams underinvest in all three layers until they have run 50+ fine-tuning experiments and cannot reproduce the one that worked best from last month. The fixes are not expensive. The cost of not having them compounds with every experiment cycle. This guide covers specific tooling choices, configurations, and real cost data from ClusterBid buyers running fine-tuning campaigns in mid-2026.

02

Dataset Versioning: DVC, HuggingFace Datasets, and Reproducibility

Dataset versioning is the foundation of reproducible fine-tuning. Without it, the model that achieved 72.3% on the evaluation benchmark last week cannot be reconstructed because the training dataset has been modified, the validation split has shifted, or the preprocessing pipeline parameters have drifted. The two most common tools in mid-2026 are DVC (Data Version Control, now at v3.4) and HuggingFace Datasets (v3.0+), which serve different use cases and are frequently used together.

DVC tracks datasets as Git-tracked pointers stored alongside the code. A dataset file version is represented as a 64-character MD5 hash in a `.dvc` file. The actual data is stored in a remote cache (S3, GCS, or an NFS path configured in `.dvc/config`). DVC operations (`dvc add`, `dvc push`, `dvc checkout`, `dvc diff`) are Git-like but operate on data rather than code. The key feature for fine-tuning is `dvc diff` between experiment branches, which shows exactly which training examples changed between run A and run B. DVC's primary limitation is that it does not natively handle structured dataset transformations (tokenization, filtering, splitting). Teams using DVC typically store raw data in DVC and apply transformations via tracked code in the DVC pipeline, which ensures reproducibility but adds a full DVC pipeline execution step per dataset version.

HuggingFace Datasets provides built-in versioning through its `Dataset.save_to_disk()` and `load_from_disk()` methods, and through HuggingFace Hub dataset versioning. A dataset pushed to the Hub with `dataset.push_to_hub('org/dataset-name', revision='v3-instruct-10k')` creates a snapshot that any team member can load with `load_dataset('org/dataset-name', revision='v3-instruct-10k')`. The HuggingFace Datasets library also tracks preprocessing transforms natively, so loading a dataset revision automatically applies the same tokenization and filtering steps that were applied during the original creation. In practice, most fine-tuning teams we work with use HuggingFace Datasets for the active dataset repository and DVC for long-term cold storage of raw data across multiple model projects, using the Hub's revision mechanism for per-experiment reproducibility and DVC's cache for cross-project data reuse.

ToolPrimary UseStorage BackendData Transform TrackingCost
DVC v3.4Raw data versioningS3, GCS, NFSVia DVC pipeline (.dvc)Free (OSS)
HuggingFace DatasetsActive training datasetsHub or local diskBuilt-in (transform map)Free (Hub: free for public)
LFS (Git LFS)Small dataset versioningGit LFS serverNone (manual)Free tier, ~$5/TB/mo
Custom (S3 + manifest)Enterprise controlS3 + metadata DBCustomS3 storage cost only
03

Experiment Tracking: WandB, MLflow, and the Cost of Metadata

Experiment tracking records hyperparameters, metrics, system metrics (GPU utilization, memory, temperature), and model artifacts for each training run. Without it, the best-performing checkpoint from a 32-run hyperparameter sweep is indistinguishable from the others after a week. WandB (Weights & Biases) and MLflow are the two dominant platforms in mid-2026, with fundamentally different cost and deployment models.

WandB is a SaaS platform with a generous free tier (100 GB of artifact storage, unlimited runs for teams up to 3 members) and paid tiers starting at $50 per user per month for Teams ($200/reporting user) and $100 per user per month for Enterprise with self-hosted options. The value proposition is convenience: one `wandb.init()` call captures GPU metrics, hyperparameters, output metrics, and model graphs, all viewable in a web dashboard with search, comparison views, and parallel coordinates plots for hyperparameter analysis. The hidden cost is artifact storage growth. A 32-run hyperparameter sweep on Llama 3.1 70B with LoRA adapters stored after each run (approximately 800 MB per adapter checkpoint) consumes 25.6 GB of artifact storage. At $0.10 per GB per month over the free tier, a team running 4 such sweeps per month accrues roughly $81 per month in artifact storage charges after the first 100 GB free tier is exhausted. The alternative is to store only the final best checkpoint in WandB and upload intermediate checkpoints to cheaper S3 storage.

MLflow (`mlflow.org`, Apache 2.0, now at v2.16) is the open-source alternative. It provides the same tracking server, model registry, and artifact storage capabilities but requires self-hosting. A typical MLflow deployment on a single 8-core VM with 32 GB RAM and 500 GB SSD (roughly $85-120 per month on AWS or $45-70 on a bare metal provider) handles tracking for a team of 10-15 ML engineers running 100-200 training runs per week. The operational overhead is the MLflow server configuration (databricks MLflow or the community distribution), backend database (PostgreSQL is standard for production), and artifact storage backend (S3 or NFS). Teams that already have a PostgreSQL instance and S3 bucket can deploy MLflow in roughly 2-4 hours and pay roughly $50-100 per month for infrastructure. The trade-off is that MLflow's UI is less polished than WandB's, particularly for hyperparameter sweep visualization and custom dashboard creation, though the gap has narrowed significantly in the 2.x releases.

PlatformModelCost at 10 Users / 200 Runs/moSelf-Host OptionSweep Visualization
WandB TeamsSaaS$500/mo + $39 artf. overage est.Enterprise onlyExcellent
WandB EnterpriseSaaS + self-host$1,000+/moYes (on-prem)Excellent
MLflow OSSSelf-host~$70/mo infraYes (always)Good (v2.16+)
TensorBoard + CometMLHybrid$0 (TB) + optional CometYes (TB)Basic
04

GPU Orchestration for Fine-Tuning: Sizing Jobs, Queues, and Co-Location

Fine-tuning jobs are fundamentally different from pre-training jobs in their GPU resource requirements. A pre-training job on Llama 3.1 405B uses 512-1024 GPUs for weeks. A fine-tuning job on the same model using LoRA uses 1-8 GPUs for hours or days. The orchestration layer must handle high job churn (potentially 50-200 discrete fine-tuning jobs per day from a team of 10 researchers), heterogeneous GPU requirements (some jobs need 1 GPU, some need 8, some need H100 for FP8 training while others can use L40S for INT4 fine-tuning), and efficient packing of multiple small jobs onto the same GPU node.

The co-location problem is the most impactful optimization for fine-tuning clusters. A single H100/node with 80 GB of VRAM running a LoRA fine-tuning on Llama 3.1 8B uses roughly 24-28 GB of VRAM (model weights at BF16 plus LoRA adapters plus optimizer states plus activations). The remaining 52-56 GB of VRAM is idle. Using a framework like FlexLoRA or the multi-adapter inference in PEFT v0.14, 2-3 independent LoRA fine-tuning jobs can co-exist on the same GPU with minimal throughput degradation (typically 5-12% per-job throughput loss from memory bandwidth contention). The effective GPU-hour cost per job drops by 55-70% because the cluster is running 2-3 jobs per GPU-hour instead of 1.

The standard mid-2026 orchestration stack for fine-tuning is Kubernetes with the NVIDIA GPU Operator, combined with a queue system like Volcano or Kueue for batch job scheduling, and PEFT with FlexLoRA or Unsloth for multi-adapter training on shared GPUs. The total infrastructure cost for a fine-tuning cluster serving a team of 15 researchers running 30-50 jobs per day is roughly 4-8 H100 GPUs at $2.85/hr, or $8,200-16,400 per month on the ClusterBid marketplace, plus approximately $1,000-1,500 per month for the control plane, storage, and observability infrastructure. That is roughly 10-15% of what the same team would spend running each job on a dedicated GPU instance, purely from the co-location multiplier.

GPU Allocation ModelJobs per GPU per Day$/JobGPU UtilizationCluster Cost/Month (8x H100)
Dedicated (1 job per GPU)2-3$2.85/hr25-35%$16,416
Co-located (2-3 LoRA jobs)6-8$0.95-1.42/hr65-80%$8,208-10,944
Mixed (LoRA + batch inference)8-12$0.71-1.07/hr75-85%$6,566-9,850
05

Distributed Training Configurations for LoRA and QLoRA

LoRA fine-tuning is not automatically distributed. The default PEFT + Transformers training loop runs on a single GPU by default. Distributed training for fine-tuning (multi-GPU or multi-node) requires explicit configuration of FSDP (Fully Sharded Data Parallel) or DeepSpeed ZeRO to shard the base model, optimizer states, and gradients across GPUs. The right choice depends on model size and GPU count.

For models up to 13B parameters on a single node with 8 GPUs, FSDP (PyTorch 2.5+'s `fully_shard` API or the higher-level `FullyShardedDataParallel` wrapper) is the standard approach. Each GPU holds 1/8 of the model weights, optimizer states, and gradients. LoRA adapter parameters (typically 0.1-1% of the model's total parameters) are replicated on every GPU, meaning the activation memory per GPU for LoRA weights is essentially zero. A fine-tuning run on Llama 3.1 8B using FSDP across 8 H100s completes roughly 6-7x faster than on a single H100 (accounting for communication overhead), and the per-GPU memory requirement drops from 28 GB to roughly 8-10 GB.

For models from 13B to 70B, QLoRA (NormalFloat4 quantization of the base model with LoRA adapters on the quantized weights) is the dominant approach. QLoRA loads the base model in 4-bit NF4 format using the BitsAndBytes library, reducing the memory requirement from 140 GB (FP16 70B) to roughly 38-42 GB. With FSDP and 4-bit base models, a 70B model can be fine-tuned on 4 H100s (with FSDP shard_degree=4) using roughly 12 GB of VRAM per GPU for the base model plus approximately 2-3 GB for LoRA weights and optimizer states. At ClusterBid mid-2026 pricing, that 4x H100 run costs $11.40 per hour, and a typical 5-hour fine-tuning run on 10,000 instruction examples costs $57. The same run without QLoRA would require 8 A100 80GB GPUs at roughly the same per-hour cost but also suffers from higher communication overhead across more GPUs.

07

Putting It Together: A Production Fine-Tuning Pipeline Stack in 2026

A production fine-tuning pipeline at mid-2026 pricing and tooling stacks the three layers together. The dataset layer: store raw data in S3 with DVC versioning. Use HuggingFace Datasets to load, preprocess, and split data by revision hash. Track every dataset push to the Hub with a branch or tag matching the experiment name. The experiment layer: run MLflow Tracking Server on a $70/month VM with PostgreSQL backend and S3 artifact store. Log hyperparameters, metrics, and dataset revision ID for every trial. Use Optuna for hyperparameter search with the MLflow callback for automatic logging.

The orchestration layer: run a Kubernetes cluster with the NVIDIA GPU Operator and Volcano queue system. Use PEFT + Unsloth with FSDP for distributed QLoRA fine-tuning on 4-8 H100s from the ClusterBid marketplace. Co-locate 2-3 LoRA jobs per GPU using FlexLoRA's memory planner. Use Kueue's cohort scheduling to pack fine-tuning jobs into available GPU capacity alongside inference or batch processing jobs. Set up Prometheus + Grafana dashboards for GPU utilization by namespace, with budget alerts that fire when any team's spend exceeds 120% of its allocation.

The total monthly infrastructure cost for a production fine-tuning pipeline supporting a team of 10-15 ML engineers running 30-40 experiments per week breaks down as follows: GPU compute (8x H100 reserved on ClusterBid at $2.50/hr with commitment) is roughly $14,400 per month. The control plane and observability stack (K8s control plane, Prometheus, MLflow, PostgreSQL) is approximately $700 per month. Dataset storage (S3 + Hub storage for 10 TB of raw and processed data) is about $250 per month. Total: roughly $15,350 per month. A team of this size running individual cloud instances on-demand would pay $30,000-45,000 per month for equivalent throughput, which is exactly the gap that proper fine-tuning infrastructure fills. For further reading on GPU procurement, see our GPU procurement strategy guide and the 2026 reserved contract playbook.

Filed under
LLM fine-tuningDataset versioningDVCHuggingFace DatasetsWandBMLflowLoRAGPU orchestration