WARM-START TRAINING: THE CONTINUOUS PRETRAINING ADVANTAGE
Continuous pretraining starts from the final checkpoint of a previous training run rather than random initialization. For a 70B model, a full pretraining run costs $2-5 million in GPU compute. A continuous pretraining update with 100B new tokens requires only 5-10 percent of the original compute: approximately $100,000-500,000. The savings come from the model already having learned syntax, reasoning, and general knowledge - the new data only needs to update factual knowledge and distributional shifts in the training corpus.
The warm-start advantage depends on the gap between training cutoff and target knowledge. For quarterly knowledge refreshes (3 months of new data), the model has 80-90 percent knowledge overlap with the existing checkpoint. For multi-year gaps (e.g., updating a 2023-cutoff model with 2024-2025 data), overlap drops to 60-70 percent, requiring more training tokens and higher learning rates. The recommended strategy: 5-20 billion tokens of new data for quarterly refreshes, 50-100 billion for annual major updates.
| Knowledge Gap | New Tokens | GPU Budget 7B | GPU Budget 70B | GPU Budget 405B | LR Schedule |
|---|---|---|---|---|---|
| 1 month (incremental) | 5B | $500-1,000 | $5,000-10,000 | $30,000-60,000 | LR=1e-5 cosine |
| 3 months (quarterly) | 20B | $2,000-4,000 | $20,000-40,000 | $120,000-240,000 | LR=3e-5 cosine |
| 1 year (major refresh) | 100B | $10,000-20,000 | $100,000-200,000 | $600K-1.2M | LR=1e-4 cosine |
| 2+ years (full reset) | 500B | $50,000-100,000 | $500K-1M | $3M-6M | LR=3e-4 warmup |
DATA MIXING: BALANCING NEW KNOWLEDGE WITH OLD DISTRIBUTION
The critical design decision in continuous pretraining is the ratio of new data to old data in each training batch. Using 100 percent new data causes catastrophic forgetting of earlier knowledge - the model can lose 5-15 percent on MMLU and other benchmarks from a single continuous training run. The standard practice mixes new and old data at ratios between 1:4 and 1:1 (new:old), with the exact ratio determined by the model scale and gap size.
For a quarterly 20B-token update on a 70B model, the optimal mix is 1:3 (5B new + 15B old per epoch, 20B total effective unique tokens). The old data is re-sampled from the original pretraining corpus, with emphasis on domains where forgetting is observed (typically code, math, and reasoning benchmarks). Processing 15B old tokens adds approximately 30-40 percent to the data loading and preprocessing pipeline but is essential for stability. The alternative approach uses replay buffers of 5-10B tokens from the original training set, updated each epoch with importance sampling based on the current model's loss.
Data deduplication between new and old corpora is a hidden GPU cost. Up to 30 percent of web-crawled data from 2024-2025 is near-duplicate of 2023 content (same events re-reported, mirrored Wikipedia updates, rehosted code repositories). Running MinHashLSH deduplication on a 100B-token corpus costs $2,000-5,000 in preprocessing compute but prevents up to 80 percent of redundant training, effectively reducing the required training budget by 20-30 percent.
| New:Old Ratio | Forgetting (MMLU) | New Knowledge Gain | Training Efficiency | Use Case |
|---|---|---|---|---|
| 1:0 (new only) | -8 to -15% | +12-18% | 100% | Domain adaptation |
| 1:1 (balanced) | -2 to -5% | +8-12% | 50% | Monthly updates |
| 1:3 (conservative) | -0.5 to -2% | +5-8% | 25% | Quarterly updates |
| 1:7 (minimal) | < -0.5% | +2-4% | 12.5% | Safety patches |
| Inverse (old heavy) | 0% | +1-2% | 10% | Stability critical |
LEARNING RATE SCHEDULES: THE WARM-UP AND DECAY PROBLEM
Continuous pretraining learning rate schedules differ fundamentally from scratch training. The standard approach uses a short warm-up (100-200 steps, versus 2,000-5,000 for scratch training) to a lower peak LR (1e-5 to 5e-5, versus 1e-4 to 3e-4 for scratch), followed by cosine decay to near-zero over the remaining steps. The reduced warm-up assumes the optimizer states from the previous run are a reasonable starting point; full re-initialization of optimizer states at the beginning of continuous training adds 5-10 percent more steps to the warm-up phase.
The scheduler interacts with GPU infrastructure through checkpointing frequency. During LR warm-up, the effective learning rate changes every step, requiring a checkpoint every 50-100 steps to capture the optimal LR peak - the peak LR rarely coincides with a round-number step boundary. Each checkpoint of a 70B model in FP16 requires 140 GB of storage and takes 30-60 seconds to write to NVMe. With 100 checkpoints across the LR ramp, this adds 1.5-2 hours to total training time and 14 TB of storage. Balancing checkpoint frequency with storage cost suggests saving optimizer states every 200 steps during warm-up and every 1,000 steps during stable training.
EVALUATION INFRASTRUCTURE FOR CONTINUOUS TRAINING
Continuous pretraining requires continuous evaluation to detect forgetting early. A standard eval suite (MMLU, HumanEval, GSM8K, HellaSwag, plus domain-specific benchmarks) on a 70B model costs $200-500 per full evaluation run in GPU compute. For 50 evaluations during a 50B-token run, evaluation adds $10,000-25,000 to the total budget - a significant 5-10 percent overhead that is often overlooked.
The solution is a tiered evaluation infrastructure: light evaluation (perplexity on 10K held-out tokens, 100 samples of MMLU) every 500 steps ($5-10 per eval), medium evaluation (full MMLU, GSM8K subsets) every 2,000 steps ($50-100), and full evaluation every 10,000 steps ($200-500). This tiered approach reduces evaluation GPU cost by 60-70 percent while maintaining detection coverage for forgetting events.
Automated rollback is essential infrastructure. If a forgetting metric drops below a threshold (e.g., MMLU drops > 2 percent from peak), the training pipeline should automatically roll back to the previous checkpoint and adjust the data mixing ratio or learning rate. The rollback infrastructure must restore the model weights, optimizer states (280 GB for 8-bit Adam on 70B), and data loader state from a previous checkpoint. NFS-based checkpoint storage with 10 Gbps networking enables rollback within 2-5 minutes.
| Eval Tier | Frequency | GPU Cost/Eval | Total Cost (50K steps) | Detection Coverage |
|---|---|---|---|---|
| Light (ppl + subset) | Every 500 steps | $5-10 | $500-1,000 | 50-60% |
| Medium (full subsets) | Every 2,000 steps | $50-100 | $1,250-2,500 | 75-85% |
| Full suite | Every 10,000 steps | $200-500 | $1,000-2,500 | 95-100% |
| Domain specific | Every 500 steps | $20-50 | $2,000-5,000 | Per-domain 90% |
| Combined tiered | Mixed frequency | $5-500 | $4,750-11,000 | 95-100% |
HARDWARE IMPLICATIONS AND INFRASTRUCTURE REQUIREMENTS
Continuous pretraining places unique demands on training infrastructure compared to scratch training. The data pipeline must handle the mixing of old and new corpora with dynamic ratios - the data loader needs to read from two separate datasets and sample at configured proportions. Each training step requires loading from multiple data sources, increasing I/O pressure by 2-3x. DAOS or GPUDirect Storage with 4+ NVMe SSDs per node in RAID-0 is recommended for 100B+ token continuous runs.
The checkpointing system must be optimized for quick rollback. Standard practice: maintain the 3 most recent checkpoints locally (on NVMe, 3 x 140 GB = 420 GB per 70B checkpoint), plus one checkpoint per 10K steps on NFS for the entire run. Total storage for a 50B continuous run: approximately 8-12 TB. The checkpoint format uses PyTorch's torch.save with _use_new_zipfile_serialization=False for compatibility with DeepSpeed's ZeRO checkpoint format, which stores one file per GPU.
B200: ECONOMICAL QUARTERLY KNOWLEDGE REFRESHES
B200's 2x memory enables continuous pretraining on larger batch sizes, directly improving the efficiency of knowledge updates. A 70B model on B200 with batch size 16 (versus 8 on H100) processes new tokens at 1.6-1.8x the throughput of H100. For a quarterly 20B-token update, training time drops from 4-6 days on 8 H100 nodes to 2-3 days on 8 B200 nodes, with similar GPU-hour cost but faster iteration cycles.
The more significant advantage is the ability to keep the full 7B model plus old-data replay buffer in GPU-accessible memory. With 192 GB VRAM, approximately 60 GB can be allocated to caching frequently replayed old-data embeddings, reducing the data loading bottleneck by 40-50 percent. For 405B-class models, B200's capacity reduces the GPU count from 128+ H100 to 64-80 B200 for equivalent throughput, making quarterly knowledge refreshes economically viable for organizations previously limited to annual updates.
