One-Off vs Continuous: Why the Infrastructure Decision Changes Completely
Most fine-tuning guides describe a one-time job: collect data, train for a few days, deploy the checkpoint, move on. Continuous fine-tuning GPU cost in 2026 is a different problem entirely. You are running a recurring job on a fixed schedule - nightly, weekly, or triggered by feedback volume thresholds - and that job needs to complete reliably, hand off a valid checkpoint, and clear capacity for inference. Every infrastructure decision that works fine for a one-time run falls apart when you repeat it 52 times a year.
The clearest signal that your team has crossed into continuous fine-tuning territory is when model quality becomes a moving target tied to production feedback. Teams doing SFT on user corrections, RLHF on A/B test winner responses, or domain adaptation on accumulating annotation batches are running production RLHF infrastructure whether they have budgeted for it or not. A 7B assistant model trained weekly on 50,000 new interaction samples is not a training project - it is a pipeline with checkpointing, evaluation gates, shadow deployment, and rollback capabilities. The GPU budget behind it needs to reflect that.
The failure mode teams hit consistently is treating continuous fine-tuning as a series of independent one-off runs. They borrow a spare eight-GPU node, use spot pricing, manage checkpoints manually. That works twice. By week three, the spot instance is unavailable at the scheduled time, the week-two checkpoint got overwritten when someone moved data to free up storage for inference, and the whole pipeline is blocked waiting for a human to intervene. Production RLHF infrastructure requires accepting that this is a steady-state workload - not a research experiment that happens to recur.
Why Spot Fails for Predictable Weekly Pipelines - and What Reserved Actually Costs
Spot GPU instances are the right answer for fault-tolerant, checkpoint-resumable, long-horizon training runs. They are the wrong answer for a weekly fine-tuning pipeline with a hard completion window. When your fine-tuning job must finish every Tuesday at 2 AM so the checkpoint can clear evaluation by Wednesday and deploy before Thursday's release, spot interruption rates of 5-15% on H100 clusters stop being a cost nuisance and become a reliability problem. Missing one window does not mean you wasted a few GPU-hours - it means your production model is four days stale and the feedback loop that drives product quality is broken.
Reserved GPU capacity for continuous fine-tuning rarely means reserving a large cluster. For a weekly SFT job on a 7B model with LoRA, you need two to four H100 SXM5 GPUs for six to twelve hours per week. Monthly H100 reserved rates from bare metal providers in 2026 run approximately $1.80-$2.20/GPU/hr on one-month terms. Spot availability for the same hardware runs $1.03-$1.50/hr when available. The 40-50% premium for reserved capacity is easy to justify when pipeline reliability goes from roughly 85% to 99.5%. At $2.00/GPU/hr reserved versus $1.25/hr spot, the cost difference for a weekly 7B LoRA run is about $60/month - trivial relative to the engineering time spent debugging a failed pipeline.
The smarter model for daily fine-tuning workloads is a hybrid: reserve a small cluster at a fixed base capacity and use spot for burst evaluation jobs. If you are training on a 70B model with full fine-tuning and need 32 GPUs for 16 hours per week, reserving that full cluster runs roughly $50,688/month at $2.20/GPU/hr. The right approach is to reserve eight GPUs for the critical path fine-tuning run and acquire the remaining 24 on spot when needed for evaluation and shadow testing. This cuts reserved spend by 75% while keeping the reliability guarantee on the part of the pipeline that cannot tolerate interruption.
The RLHF Production Stack: How Actor, Reference, and Reward Models Split Your GPU Budget
Production RLHF - specifically PPO and its newer variants GRPO (Group Relative Policy Optimization) and DPO (Direct Preference Optimization) - requires a fundamentally different cluster topology than supervised fine-tuning. In PPO, you run four distinct compute-intensive processes simultaneously: the actor model (the policy being trained), the reference model (the frozen base), the critic/value model, and the reward model. Each needs GPU memory. The naive implementation loads all four on separate GPUs, which makes the memory math quickly alarming.
For a 7B PPO setup, the practical minimum is eight H100 SXM5 GPUs: two for the actor (trainable, needs gradient state at roughly 3x model weights in BF16 with Adam), two for the reference model (inference-only, frozen), two for the critic, and two for the reward model. With LoRA applied to the actor and critic, you can reduce this to four H100s, but throughput drops significantly. Teams at mid-scale production use sixteen to thirty-two H100s for 7B PPO to get useful tokens-per-second rates. For 70B PPO, expect 64-128 H100 SXM5 GPUs for a realistic production setup. The actor alone at BF16 with ZeRO-3 sharding across eight GPUs consumes most of the available HBM budget before the other three model copies are even loaded.
GRPO has gained significant adoption in 2026 for production RLHF because it removes the critic model from the training loop entirely. This reduces GPU memory per training step by 20-30% and eliminates one of the four concurrent model copies. For 70B GRPO in production, you can run on 32-64 H100 SXM5 GPUs instead of 64-128 for full PPO - a meaningful infrastructure difference. DPO eliminates both the critic and the reward model, reducing the cluster to just the policy and reference: GPU requirements drop to roughly half of PPO for equivalent model sizes. DPO cannot use online feedback signals, limiting it to datasets labeled ahead of time, but for teams with a human preference labeling pipeline it is the most cost-effective continuous RLHF option.
| Algorithm | 7B GPU Requirement | 70B GPU Requirement |
|---|---|---|
| PPO (full, production) | 16-32x H100 SXM5 | 64-128x H100 SXM5 |
| PPO (with LoRA) | 4-8x H100 SXM5 | 16-32x H100 SXM5 |
| GRPO (no critic) | 8-16x H100 SXM5 | 32-64x H100 SXM5 |
| DPO (offline only) | 4-8x H100 SXM5 | 16-32x H100 SXM5 |
Daily SFT Cost Model: GPU-Hours per Day for 7B, 13B, and 70B Models
Supervised fine-tuning on production feedback data is the most common form of continuous fine-tuning GPU cost in 2026. How many GPU-hours you consume per run depends on three variables: dataset size, model size, and whether you are doing full fine-tuning or parameter-efficient fine-tuning with LoRA or QLoRA. The numbers that follow assume a daily fine-tuning cadence on a dataset of 50,000 new training samples, each averaging 512 tokens - a realistic volume for a mid-scale AI product accumulating user feedback daily.
For a 7B model with QLoRA (rank 16, targeting attention and MLP projections), a daily run on 50,000 samples takes approximately 1.5-2.5 GPU-hours on a single H100 SXM5. At $1.80/hr reserved, that is $2.70-$4.50/day. Full fine-tuning the same 7B model on the same dataset takes 6-10 GPU-hours on two H100 SXM5 GPUs with gradient checkpointing and ZeRO-2, or $21.60-$36.00/day. The 8-10x cost difference explains why teams choose LoRA by default, but there is a real reason to choose full fine-tuning for continuous pipelines: after 30-50 LoRA cycles on the same base, adapter weights accumulate drift that requires periodic remerging and retraining from the base. Full fine-tuning sidesteps this at higher cost per run.
The 70B model numbers are where the conversation shifts. A 70B QLoRA run (rank 8, attention layers only) on 50,000 samples takes approximately 8-14 GPU-hours on four H100 SXM5 GPUs at BF16. That is $57.60-$100.80/day at reserved rates. Full fine-tuning a 70B model daily is the decision that requires board-level GPU budget approval: expect 64-96 GPU-hours on eight H100 SXM5 GPUs, translating to $460-$691/day, or $14,000-$21,000/month. Almost every team running a 70B model shifts to weekly full fine-tuning plus daily LoRA adaptation layers on top, merging the adapters into the base checkpoint weekly. This hybrid runs the fine-tuning budget around $1,500-$3,000/month - manageable for any company with a production AI product.
| Configuration | GPU-Hours/Day | Daily Cost (H100 Reserved) |
|---|---|---|
| 7B QLoRA (50K samples) | 1.5-2.5 GPU-h | $2.70-$4.50 |
| 7B Full FT (50K samples) | 6-10 GPU-h | $21.60-$36.00 |
| 13B QLoRA (50K samples) | 3-5 GPU-h | $10.80-$18.00 |
| 13B Full FT (50K samples) | 14-22 GPU-h | $50.40-$79.20 |
| 70B QLoRA (50K samples) | 8-14 GPU-h | $57.60-$100.80 |
| 70B Full FT (50K samples) | 64-96 GPU-h | $460-$691 |
Multi-Cluster Architecture: Why Fine-Tuning and Inference Cannot Share Capacity
The most common failure mode in production RLHF infrastructure is capacity contention. Teams that run fine-tuning jobs on the same GPU cluster serving production inference discover this problem at the worst moment: a fine-tuning job that kicks off at 2 AM adds latency to live inference requests, triggers OOM errors when competing with the KV cache for HBM, and gets preempted mid-checkpoint when inference autoscaler sees traffic spike. The fix is not better scheduling. It is separate clusters. Running separate GPU pools for fine-tuning and inference is the right architecture, and it is less expensive than it sounds.
The fine-tuning cluster does not need to be online 24 hours a day. For weekly fine-tuning, a reserved cluster of four to eight H100 SXM5 GPUs runs for 8-16 hours per week. The rest of the time those GPUs sit idle unless you fill them with other training workloads. This is where fine-tuning procurement differs from inference: inference clusters need 100% availability and single-digit millisecond response times. Fine-tuning clusters can tolerate pre-provisioning delays of a few minutes and need only throughput. The practical implication is that fine-tuning capacity should be sourced differently - bare metal reserved capacity rather than cloud GPU instances with VM startup latency. Fine-tuning JCT (job completion time) variance on cloud VMs is typically 30-50% higher than bare metal due to noisy neighbors and memory bandwidth sharing.
Clean cluster separation also makes proper model versioning possible. When fine-tuning completes, the new checkpoint goes through evaluation (1-4 GPU-hours on a small evaluation cluster for a 7B model, 8-16 GPU-hours for 70B on a comprehensive benchmark suite), then shadow traffic testing before replacing the production model. This pipeline requires a place to run the new model in shadow while the current production model handles live traffic. That is impossible if fine-tuning and inference share the same hardware. Teams that separate their production RLHF infrastructure from inference consistently report faster deployment cadences and better confidence in each model update - because the deployment pipeline is defined and repeatable, not improvised around shared capacity.
GRPO vs DPO vs PPO: How Algorithm Choice Shapes Your Cluster More Than Model Size
The algorithm you choose for production RLHF is the single largest determinant of cluster size and cost - larger than model size in many configurations. Full PPO at 70B scale with four concurrent model copies needs 64-128 H100 SXM5 GPUs. At $2.20/GPU/hr reserved for the top of that range, you are committing $202,752/month to your continuous fine-tuning infrastructure. Very few companies can justify that on a permanent basis. The algorithm selection conversation is therefore not just a machine learning question - it is the central GPU procurement decision for any team doing production RLHF.
GRPO deserves serious consideration for teams currently running full PPO. It achieves comparable results on instruction-following and reasoning benchmarks by using group-relative reward normalization instead of a learned critic. The critic removal drops the cluster requirement to 32-64 H100 SXM5 GPUs for 70B GRPO - a 50% reduction versus full PPO. Teams migrating from PPO to GRPO in 2026 have reported minimal quality regression on most production tasks, with the main exception being complex multi-step reasoning chains where the critic's learned value function provides signal the group-relative approach misses. If your task is instruction following, summarization, or code generation rather than long-horizon reasoning, GRPO is likely the right move.
DPO is the most GPU-efficient RLHF approach available, and it is underused in production. For teams with a human preference labeling pipeline producing 5,000-20,000 labeled pairs per week, weekly DPO fine-tuning on a 70B model needs only 16-32 H100 SXM5 GPUs for 12-24 hours, which at reserved rates runs roughly $8,448-$16,896/month. The limitation is hard: DPO cannot incorporate automated reward signals (code execution results, mathematical verification, factual grounding scores). If your quality signal comes from automated evaluation rather than human labeling, DPO is not viable. But if you are running an AI product where human preference labeling is feasible, DPO on a small reserved cluster is among the most cost-effective forms of continuous fine-tuning available.
Capacity Planning for Steady-State Fine-Tuning: Building the Weekly Budget Model
Continuous fine-tuning GPU cost should appear as a recurring line item in your AI infrastructure budget, not a project expense. A team running weekly SFT on a 7B model with QLoRA spends roughly $200-$400/month on GPU time at reserved H100 rates. Weekly full fine-tuning on a 13B model runs $800-$1,500/month. At this cost level, the conversation is less about GPU cost and more about operational reliability - whether the pipeline runs consistently every week and whether checkpoint management, evaluation, and deployment are automated rather than manual.
For teams running 70B models with RLHF, capacity planning becomes more significant. Weekly GRPO on a 70B model needs approximately 32 H100 SXM5 GPUs for 12-18 hours per week. At $2.20/GPU/hr reserved, that is $16,896-$25,344/month. Add evaluation cluster overhead (8-16 GPU-hours per checkpoint) and shadow deployment capacity, and total fine-tuning infrastructure for a production 70B model typically runs $20,000-$35,000/month. This is the level where multi-month reserved contracts start to matter - a six-month commitment on the right cluster from ClusterBid's sourcing desk typically yields 20-35% savings versus month-to-month rates, which compounds meaningfully on a $300,000/year infrastructure line item.
The variable most teams underestimate in their capacity plan is evaluation overhead. Every fine-tuning run that produces a new checkpoint requires evaluation before deployment. For a 7B model evaluated on a 5,000-sample benchmark suite, plan for 2-4 GPU-hours on a single H100. For a 70B model on a comprehensive suite covering domain benchmarks, safety evaluations, and regression tests, plan for 8-16 GPU-hours per checkpoint. If you run evaluation concurrently rather than sequentially with fine-tuning, you need evaluation cluster capacity in addition to fine-tuning cluster capacity. Teams that budget for evaluation GPU-hours separately from fine-tuning GPU-hours end up with much more accurate infrastructure cost models.
Sourcing Reliable Production Fine-Tuning Capacity: What to Look for in a Contract
Continuous fine-tuning pipelines have a specific procurement profile that most providers struggle to serve well. You need guaranteed availability at a specific day and time each week, consistent job startup latency (not 2-minute variation that throws off your pipeline scheduling), and ideally the ability to pause reserved capacity during weeks when fine-tuning is skipped - A/B testing periods, data quality holds, model freezes. Cloud on-demand GPU instances fail the first requirement. Spot instances fail the first and second. Most shared compute platforms with dynamic scheduling fail all three.
Bare metal reserved capacity is the correct infrastructure category for production RLHF and continuous SFT workloads. You get deterministic availability, direct hardware access without hypervisor overhead, and full control over the software stack - which matters for NCCL tuning, CUDA version pinning, and the distributed training library versions that your checkpointing code depends on. For fine-tuning specifically, avoiding VM overhead produces more consistent JCT than any scheduling optimization you can apply at the software layer. The thing nobody tells you about cloud GPU VMs for training is that noisy neighbor variance affects not just average throughput but checkpoint timing - a training step that takes 820ms in isolation takes 900-1100ms on a loaded hypervisor, which compounds over millions of steps.
The most efficient procurement model for continuous fine-tuning is a short-term reserved contract sized for your base fine-tuning workload (four to eight GPUs for most 7B-13B pipelines, eight to thirty-two for 70B) plus access to on-demand capacity for evaluation burst and occasional larger runs. ClusterBid's inventory includes H100 SXM5 and H200 SXM clusters from vetted bare metal providers on terms starting at one month, sourced competitively across multiple providers to get pricing below what you can access through direct provider relationships. If your continuous fine-tuning pipeline is a permanent fixture of your model quality strategy rather than an experiment, the reserved contract desk is the right place to start.
