DISTILLATION PARADIGMS AND THEIR GPU PROFILES
Knowledge distillation transfers capabilities from a large teacher model to a smaller student model. Three paradigms dominate: offline distillation where the teacher generates a static dataset of logits/features for the student to train on; online distillation where the teacher and student train simultaneously with the teacher generating soft targets for each student batch; and on-policy distillation where the student generates outputs and the teacher scores or corrects them. Each has a fundamentally different GPU infrastructure profile.
Offline distillation has the highest up-front GPU cost: generating scores for the teacher on the full training dataset. For a 405B teacher generating soft targets for 1 trillion tokens, the teacher inference run costs $1-3 million in GPU compute. But once generated, the student training proceeds independently, often on 80 percent fewer GPUs than the teacher required. Online distillation has lower peak cost but ties teacher and student GPUs together for the full training duration, increasing the training cluster size by 2-4x. On-policy distillation is the most GPU-intensive because every student generation requires teacher evaluation, similar to RLHF's four-model setup but with typically larger teacher-student size differences.
| Paradigm | Teacher GPU Profile | Student GPU Profile | Total GPU Cost | Quality Transfer | Best For |
|---|---|---|---|---|---|
| Offline (logit) | High pre-training, zero during training | Standard training | $1-3M (teacher) + $100-500K (student) | 90-95% | 1T token scale |
| Offline (feature) | High pre-training | Standard training + feature loss | $500K-2M + $100-400K | 85-92% | Vision models |
| Online (logit) | Live inference throughout training | Standard + teacher forward | $2-5M total (200K steps) | 93-97% | Medium scale |
| On-policy (RL) | Scoring student outputs | Generation + training | $3-8M total | 95-98% | Reasoning models |
TEACHER INFERENCE SERVING: THE BOTTLENECK
In offline distillation, the teacher inference run requires serving the teacher model at maximum throughput - not latency. For a 405B teacher generating soft targets on 1T tokens (approximately 500B teacher-forward passes), the compute requirement is approximately 5-15 exaFLOPs. On H100 clusters at 40 percent utilization, this takes 15,000-50,000 GPU-hours. At $3/GPU-hour, the teacher-serving budget is $45,000-150,000 - a significant line item but 10-30x cheaper than training a 405B model from scratch.
Teacher serving uses maximum batch sizes (128-256) with vLLM or TensorRT-LLM optimized for throughput rather than latency. The key optimization is KV-cache management: the teacher processes each sequence exactly once (forward pass, no generation), so no KV cache retention is needed. This allows using the maximum possible batch size that fits in GPU memory for a single forward pass. For a 405B model at FP16 with batch size 128 and 4K sequence length, teacher inference requires 810 GB (weights) + 512 GB (activations) = 1.3 TB total, requiring 20-24 H100 GPUs for the teacher serving pool.
Distributed teacher serving across multiple nodes introduces a latency bottleneck. Each batch of 128 sequences at 4K tokens generates 512K tokens per forward pass. At 200K steps to cover 1T tokens (with each step processing 5M tokens), the teacher processes 5M tokens per step. On 24 H100 GPUs at approximately 2,000 tokens/second/GPU (405B FP16), each step takes 100-120 seconds. Total wall-clock time: 200K steps x 110 seconds = 230 days! This forces either smaller model quantization (FP8: 500 tokens/s/GPU, 4x fewer GPUs, 2x faster) or using a smaller teacher (70B: 5,000 tokens/s/GPU, 24 GPUs, completes in 10-20 days).
| Teacher Size | Throughput per H100 | GPUs Needed | 1T Token Time | GPU Cost | Recommendation |
|---|---|---|---|---|---|
| 405B FP16 | 350 tok/s/GPU | 64-128 | 35-70 days | $180K-360K | Too expensive |
| 405B FP8 | 700 tok/s/GPU | 32-64 | 18-35 days | $40K-180K | Good for frontier |
| 70B FP16 | 4,500 tok/s/GPU | 8-16 | 8-16 days | $3K-8K | Best cost/value |
| 70B FP8 | 8,000 tok/s/GPU | 4-8 | 4-8 days | $1K-3K | Most practical |
| 13B FP16 | 20,000 tok/s/GPU | 2-4 | 2-4 days | $300-800 | Small student only |
ONLINE DISTILLATION: TEACHER-STUDENT CO-TRAINING INFRASTRUCTURE
Online distillation runs teacher and student on synchronized GPU pools. The teacher processes each batch to generate soft targets, which are consumed by the student immediately. The teacher runs at inference throughput (no backward pass), while the student runs at training throughput (forward + backward + optimizer). The ratio of teacher GPUs to student GPUs depends on their relative sizes: for a 70B teacher and a 7B student, the teacher processes approximately 1.5-2x the tokens per GPU compared to the student's training throughput, so 1 teacher GPU serves 2-3 student GPUs.
The infrastructure design uses a teacher inference pool that pre-generates soft targets for multiple student steps in advance. A queue of pre-computed teacher logits (FP32, batch size x vocabulary size) is maintained in CPU memory. For a 7B student with vocabulary 128K and batch size 32, each batch of teacher logits consumes 32 x 128K x 4 bytes = 16 MB. A queue depth of 100 batches consumes 1.6 GB CPU memory - negligible. The teacher pool pre-fills the queue ahead of the student's consumption rate, handling transient slowdowns without synchronizing the two pipelines.
LOGIT MATCHING VERSUS FEATURE DISTILLATION: COMPUTE TRADEOFFS
Logit matching (soft target distillation) requires the full vocabulary logits from the teacher. For a 128K vocabulary, each token generates a 128K-dimensional logit vector - 512 KB per token at FP32, or 2 GB per 4K-token sequence. The I/O and memory cost of storing and streaming these logits during offline distillation is significant: 1 trillion tokens generates 1 trillion x 512 KB = 512 exabytes of soft targets - obviously infeasible. In practice, offline logit distillation uses top-k (typically k=8-16) or temperature-scaled probability distillation that stores only the top-k probabilities per token, reducing the 512 exabytes to 32-64 TB.
Feature distillation matches intermediate representations (hidden states at specific layers) rather than output logits. This reduces the per-token data from 512 KB to 4-16 KB (hidden dimension size), making offline storage feasible. However, feature distillation requires the student architecture to have the same hidden dimension at the distillation layer, constraining architectural freedom. The GPU compute cost is also higher during training because both teacher and student must compute forward passes to the distillation layer, versus logit matching which only requires teacher output computation.
The empirical conclusion: logit matching with top-16 soft targets is the most compute-efficient approach for general domain distillation, while feature distillation provides 2-5 percent better quality transfer for specific tasks (reasoning, code generation) at 30-50 percent higher training GPU cost.
| Method | Storage per Token | 1T Token Storage | Training GPU Overhead | Quality Transfer | Infrastructure |
|---|---|---|---|---|---|
| Full logit (vocab) | 512 KB | 512 EB (infeasible) | 0% (pre-computed) | 95-97% | Impossible to store |
| Top-k logit (k=16) | 8 KB | 8 TB | 0% (pre-computed) | 93-96% | Feasible offline |
| Temperature softmax | 8 KB (top-16) | 8 TB | 0% | 94-97% | Best offline |
| Hidden state (last) | 8 KB (4096-d) | 8 TB | 15-25% | 91-95% | Same arch only |
| Multi-layer feature | 64 KB (8 layers) | 64 TB | 30-50% | 93-97% | Best quality |
ON-POLICY DISTILLATION AND SYNTHETIC DATA GENERATION
DeepSeek-R1 and similar reasoning models have popularized on-policy distillation: the student generates chain-of-thought reasoning traces, and the teacher scores or corrects them. This is the most GPU-intensive paradigm because each student training step requires: student generation (inference), teacher evaluation (inference), and student training (forward + backward). For a 70B teacher and 7B student, each training step processes 3x the tokens of standard training.
The GPU cluster design for on-policy distillation separates the pipeline into a student generation pool (running 7B inference at high throughput), a teacher scoring pool (running 70B inference at moderate throughput), and a student training pool (running 7B training). With a 10:1 teacher-to-student size ratio, approximately 1 teacher GPU serves 4-6 student training GPUs and 8-10 student generation GPUs. The generation pool is typically the largest: each student training step requires 4-8 generation steps to fill a replay buffer, requiring 4-8x more generation throughput than training throughput.
B200: MAKING TEACHER INFERENCE PRACTICAL
B200's 192 GB VRAM allows a 70B teacher to serve with batch size 128-256 on a single GPU, compared to 32-64 on H100. For the teacher inference pass in offline distillation, a 4-GPU B200 node matches the throughput of an 8-GPU H100 node at similar cost. For a 405B teacher, B200 enables FP8 deployment on 8 GPUs versus 16-24 H100, halving the teacher inference budget.
The combined effect for a full offline distillation pipeline (1T tokens, 405B teacher, 7B student): H100 cluster = 32 teacher GPUs + 8 student GPUs, 40 days, $85,000. B200 cluster = 16 teacher GPUs + 4 student GPUs, 25 days, $50,000. The 40 percent cost reduction and 38 percent time reduction make quarterly distillation cycles feasible for organizations previously limited to annual distillation.
