All essays
InfrastructureINFRASTRUCTUREFEB 2026

Knowledge Distillation Infrastructure: Student-Teacher GPU Cluster Design

GPU infrastructure for knowledge distillation at scale: teacher inference serving, student training infrastructure, offline vs online distillation, logit matching vs feature distillation. GPU cluster architecture, throughput benchmarks, and cost analysis for distilling models from 405B teacher to 7B-70B students.

01

DISTILLATION PARADIGMS AND THEIR GPU PROFILES

Knowledge distillation transfers capabilities from a large teacher model to a smaller student model. Three paradigms dominate: offline distillation where the teacher generates a static dataset of logits/features for the student to train on; online distillation where the teacher and student train simultaneously with the teacher generating soft targets for each student batch; and on-policy distillation where the student generates outputs and the teacher scores or corrects them. Each has a fundamentally different GPU infrastructure profile.

Offline distillation has the highest up-front GPU cost: generating scores for the teacher on the full training dataset. For a 405B teacher generating soft targets for 1 trillion tokens, the teacher inference run costs $1-3 million in GPU compute. But once generated, the student training proceeds independently, often on 80 percent fewer GPUs than the teacher required. Online distillation has lower peak cost but ties teacher and student GPUs together for the full training duration, increasing the training cluster size by 2-4x. On-policy distillation is the most GPU-intensive because every student generation requires teacher evaluation, similar to RLHF's four-model setup but with typically larger teacher-student size differences.

ParadigmTeacher GPU ProfileStudent GPU ProfileTotal GPU CostQuality TransferBest For
Offline (logit)High pre-training, zero during trainingStandard training$1-3M (teacher) + $100-500K (student)90-95%1T token scale
Offline (feature)High pre-trainingStandard training + feature loss$500K-2M + $100-400K85-92%Vision models
Online (logit)Live inference throughout trainingStandard + teacher forward$2-5M total (200K steps)93-97%Medium scale
On-policy (RL)Scoring student outputsGeneration + training$3-8M total95-98%Reasoning models
02

TEACHER INFERENCE SERVING: THE BOTTLENECK

In offline distillation, the teacher inference run requires serving the teacher model at maximum throughput - not latency. For a 405B teacher generating soft targets on 1T tokens (approximately 500B teacher-forward passes), the compute requirement is approximately 5-15 exaFLOPs. On H100 clusters at 40 percent utilization, this takes 15,000-50,000 GPU-hours. At $3/GPU-hour, the teacher-serving budget is $45,000-150,000 - a significant line item but 10-30x cheaper than training a 405B model from scratch.

Teacher serving uses maximum batch sizes (128-256) with vLLM or TensorRT-LLM optimized for throughput rather than latency. The key optimization is KV-cache management: the teacher processes each sequence exactly once (forward pass, no generation), so no KV cache retention is needed. This allows using the maximum possible batch size that fits in GPU memory for a single forward pass. For a 405B model at FP16 with batch size 128 and 4K sequence length, teacher inference requires 810 GB (weights) + 512 GB (activations) = 1.3 TB total, requiring 20-24 H100 GPUs for the teacher serving pool.

Distributed teacher serving across multiple nodes introduces a latency bottleneck. Each batch of 128 sequences at 4K tokens generates 512K tokens per forward pass. At 200K steps to cover 1T tokens (with each step processing 5M tokens), the teacher processes 5M tokens per step. On 24 H100 GPUs at approximately 2,000 tokens/second/GPU (405B FP16), each step takes 100-120 seconds. Total wall-clock time: 200K steps x 110 seconds = 230 days! This forces either smaller model quantization (FP8: 500 tokens/s/GPU, 4x fewer GPUs, 2x faster) or using a smaller teacher (70B: 5,000 tokens/s/GPU, 24 GPUs, completes in 10-20 days).

Teacher SizeThroughput per H100GPUs Needed1T Token TimeGPU CostRecommendation
405B FP16350 tok/s/GPU64-12835-70 days$180K-360KToo expensive
405B FP8700 tok/s/GPU32-6418-35 days$40K-180KGood for frontier
70B FP164,500 tok/s/GPU8-168-16 days$3K-8KBest cost/value
70B FP88,000 tok/s/GPU4-84-8 days$1K-3KMost practical
13B FP1620,000 tok/s/GPU2-42-4 days$300-800Small student only
03

ONLINE DISTILLATION: TEACHER-STUDENT CO-TRAINING INFRASTRUCTURE

Online distillation runs teacher and student on synchronized GPU pools. The teacher processes each batch to generate soft targets, which are consumed by the student immediately. The teacher runs at inference throughput (no backward pass), while the student runs at training throughput (forward + backward + optimizer). The ratio of teacher GPUs to student GPUs depends on their relative sizes: for a 70B teacher and a 7B student, the teacher processes approximately 1.5-2x the tokens per GPU compared to the student's training throughput, so 1 teacher GPU serves 2-3 student GPUs.

The infrastructure design uses a teacher inference pool that pre-generates soft targets for multiple student steps in advance. A queue of pre-computed teacher logits (FP32, batch size x vocabulary size) is maintained in CPU memory. For a 7B student with vocabulary 128K and batch size 32, each batch of teacher logits consumes 32 x 128K x 4 bytes = 16 MB. A queue depth of 100 batches consumes 1.6 GB CPU memory - negligible. The teacher pool pre-fills the queue ahead of the student's consumption rate, handling transient slowdowns without synchronizing the two pipelines.

04

LOGIT MATCHING VERSUS FEATURE DISTILLATION: COMPUTE TRADEOFFS

Logit matching (soft target distillation) requires the full vocabulary logits from the teacher. For a 128K vocabulary, each token generates a 128K-dimensional logit vector - 512 KB per token at FP32, or 2 GB per 4K-token sequence. The I/O and memory cost of storing and streaming these logits during offline distillation is significant: 1 trillion tokens generates 1 trillion x 512 KB = 512 exabytes of soft targets - obviously infeasible. In practice, offline logit distillation uses top-k (typically k=8-16) or temperature-scaled probability distillation that stores only the top-k probabilities per token, reducing the 512 exabytes to 32-64 TB.

Feature distillation matches intermediate representations (hidden states at specific layers) rather than output logits. This reduces the per-token data from 512 KB to 4-16 KB (hidden dimension size), making offline storage feasible. However, feature distillation requires the student architecture to have the same hidden dimension at the distillation layer, constraining architectural freedom. The GPU compute cost is also higher during training because both teacher and student must compute forward passes to the distillation layer, versus logit matching which only requires teacher output computation.

The empirical conclusion: logit matching with top-16 soft targets is the most compute-efficient approach for general domain distillation, while feature distillation provides 2-5 percent better quality transfer for specific tasks (reasoning, code generation) at 30-50 percent higher training GPU cost.

MethodStorage per Token1T Token StorageTraining GPU OverheadQuality TransferInfrastructure
Full logit (vocab)512 KB512 EB (infeasible)0% (pre-computed)95-97%Impossible to store
Top-k logit (k=16)8 KB8 TB0% (pre-computed)93-96%Feasible offline
Temperature softmax8 KB (top-16)8 TB0%94-97%Best offline
Hidden state (last)8 KB (4096-d)8 TB15-25%91-95%Same arch only
Multi-layer feature64 KB (8 layers)64 TB30-50%93-97%Best quality
05

ON-POLICY DISTILLATION AND SYNTHETIC DATA GENERATION

DeepSeek-R1 and similar reasoning models have popularized on-policy distillation: the student generates chain-of-thought reasoning traces, and the teacher scores or corrects them. This is the most GPU-intensive paradigm because each student training step requires: student generation (inference), teacher evaluation (inference), and student training (forward + backward). For a 70B teacher and 7B student, each training step processes 3x the tokens of standard training.

The GPU cluster design for on-policy distillation separates the pipeline into a student generation pool (running 7B inference at high throughput), a teacher scoring pool (running 70B inference at moderate throughput), and a student training pool (running 7B training). With a 10:1 teacher-to-student size ratio, approximately 1 teacher GPU serves 4-6 student training GPUs and 8-10 student generation GPUs. The generation pool is typically the largest: each student training step requires 4-8 generation steps to fill a replay buffer, requiring 4-8x more generation throughput than training throughput.

06

B200: MAKING TEACHER INFERENCE PRACTICAL

B200's 192 GB VRAM allows a 70B teacher to serve with batch size 128-256 on a single GPU, compared to 32-64 on H100. For the teacher inference pass in offline distillation, a 4-GPU B200 node matches the throughput of an 8-GPU H100 node at similar cost. For a 405B teacher, B200 enables FP8 deployment on 8 GPUs versus 16-24 H100, halving the teacher inference budget.

The combined effect for a full offline distillation pipeline (1T tokens, 405B teacher, 7B student): H100 cluster = 32 teacher GPUs + 8 student GPUs, 40 days, $85,000. B200 cluster = 16 teacher GPUs + 4 student GPUs, 25 days, $50,000. The 40 percent cost reduction and 38 percent time reduction make quarterly distillation cycles feasible for organizations previously limited to annual distillation.

Filed under
Knowledge Distillation GPUStudent Teacher TrainingModel Distillation InfrastructureOffline Distillation PipelineLogit Matching GPUFeature Distillation TrainingTeacher Inference Serving