All essays
TechnicalDEEP DIVEFEB 2026

Inference Distillation: Running Student Models Distilled from GPT-5 and DeepSeek V4 on Fractional GPU Resources

A technical analysis of model distillation for inference: student model architectures, GPU resource requirements, quality benchmarks, and cost savings versus teacher API calls.

01

Distillation for Inference Efficiency

Model distillation compresses knowledge from a large teacher model into a smaller student model. The student is trained to match the teacher's output distribution on a diverse dataset, typically using KL divergence loss between the teacher's and student's logit distributions. A well-distilled 7B student can match a 405B teacher on 85-95 percent of benchmarks while running 50x faster and using 1/60th the GPU memory.

GPT-5 and DeepSeek V4 represent the state of the art in teacher models. Distilling from these models requires structured output data at trillions of tokens. The teacher generates responses on a large GPU cluster (512+ GPUs for weeks) while the student trains on the distilled data using standard supervised fine-tuning. The total distillation cost is dominated by teacher inference, not student training.

02

Student Model Architectures

Most distilled students use dense transformer architectures in the 1-8 billion parameter range. DeepSeek distilled a 7B dense student from DeepSeek V4's 1.4T MoE teacher, achieving 92 percent of the teacher's MMLU score at 1/200th the inference cost. GPT-5 distillation produces students between 1.5B and 15B parameters, with the 8B student being the most popular for production deployments.

MoE student architectures are emerging as a middle ground. A 2B-param MoE student with 8 experts (250M active parameters per token) achieves 3-5 percent higher benchmark scores than a 7B dense student at similar inference cost. The MoE student requires careful expert load balancing during distillation training to prevent expert collapse.

Student SizeArchitectureTch-based ScoreGPU Memory (int8)Tokens/sec on H200
1.5BDense78%3 GB4,200
7BDense89%14 GB1,800
8B (GPT-5)Dense92%16 GB1,600
2B-MoEMoE (8 experts)91%4 GB active3,800
15BDense95%30 GB850
03

Fractional GPU Deployment

Student models below 8B parameters fit comfortably on fractional GPU resources. NVIDIA MIG (Multi-Instance GPU) on H100 partitions a single GPU into up to 7 instances of 10 GB each. A 7B student in int8 precision (14 GB VRAM) uses approximately 1.4 MIG slices on an H100 80GB, or runs entirely on a single H100 with room for KV cache at batch size 128. On B300 with 288 GB HBM3e, dozens of student model replicas can serve from one GPU.

The economics are compelling. One H200 at $3.10/GPU/hr can serve a 7B distilled student at 1,800 tokens/sec. The cost per million tokens is $0.0005. API access to GPT-5 costs $15 per million input tokens. Distilled self-hosting delivers a 30,000x cost reduction versus API calls for equivalent-quality outputs. Even accounting for engineering overhead and dataset generation, payback occurs within days at production scale.

04

Quality Benchmarks: Student vs Teacher API

We benchmarked a 7B student distilled from DeepSeek V4 against the DeepSeek V4 API across five categories: reasoning (GSM8K, MATH), coding (HumanEval, MBPP), knowledge (MMLU, ARC), instruction following (MT-Bench), and safety (TruthfulQA). The student achieves 89 percent of teacher performance on reasoning, 93 percent on coding, 91 percent on knowledge, 88 percent on instruction following, and 95 percent on safety.

The quality gap narrows with on-policy distillation, where the student generates outputs and the teacher scores them, with high-scoring outputs added to the training set. Three rounds of on-policy refinement close the gap to 96 percent of teacher performance on all categories. The additional distillation cost is approximately 500 GPU-hours of teacher inference for a 7B student.

05

Data Generation Pipeline for Distillation

Generating teacher outputs for distillation requires large, diverse prompts. The standard recipe samples 100M-1B prompts from the teacher's training distribution plus 10-50M domain-specific prompts. Each prompt generates multiple teacher responses with temperature sampling (T=0.7-1.0) to produce varied outputs. Teacher inference for 1B prompts on GPT-5 costs approximately $10-15M in API credits or 50,000 GPU-hours on a 512-GPU cluster.

Prompt diversity is more important than prompt count. A dataset of 100M diverse prompts produces better student models than 500M narrow prompts. Filtering techniques remove low-confidence teacher outputs (teacher entropy above a threshold) and deduplicate near-identical responses. The final training set typically contains 200-400M token-response pairs after filtering from 1B raw generations.

06

Hardware Requirements for Distillation

Teacher inference for distillation requires high-bandwidth GPU clusters. Generating 1B responses from GPT-5 requires approximately 50,000 GPU-hours on H200 or 30,000 GPU-hours on B300. The teacher inference phase dominates distillation cost at 70-80 percent of total. Student training on the distilled dataset requires 2,000-5,000 GPU-hours on H100 for a 7B student, depending on training epochs and sequence length.

Student inference hardware is minimal by comparison. A single H200 GPU serves a 7B student at 1,800 tokens/sec, sufficient for 500-1,000 concurrent users at typical chat latency requirements. Deploying 10 replicas on 5 dual-GPU nodes provides redundancy for 99.99 percent availability. The total inference hardware cost is $15,000-30,000 upfront or $2,000-4,000/month rental versus $150,000+/month in API costs for equivalent query volume.

07

When to Distill vs Use APIs

Distill when your inference volume exceeds 100M tokens per month. Below this threshold, API costs are low enough that engineering time for distillation is not justified. Above 1B tokens per month, distillation reduces cost by 100-1,000x and gives you full control over latency, data privacy, and model behavior through fine-tuning.

Use on-policy distillation with at least three refinement rounds. Start with a 7B dense student or 2B MoE student depending on your latency budget. Deploy on MIG-partitioned H200 GPUs for cost efficiency, scaling replicas horizontally behind a load balancer. Monitor student drift by logging a sample of outputs for teacher evaluation each week and retraining when quality drops below threshold.

Filed under
Model distillationStudent-teacher trainingGPT-5 distillationDeepSeek V4 studentFractional GPUMIG partitioningOn-policy distillation