All essays
TechnicalDEEP DIVEFEB 2026

AI Model Compression in 2026: Pruning, Distillation, and Quantization Techniques

Comprehensive comparison of model pruning, knowledge distillation, and quantization (FP8 INT4 NF4) for LLMs. Compression ratios GPU memory savings and accuracy trade-offs for inference deployment at mid-2026.

01

Why Compression Matters for GPU Economics

The GPU budget for inference is directly proportional to model size. A 70B-parameter LLM in FP16 requires 140 GB of GPU memory, which demands at least two H100 (80 GB each) or one B200 (192 GB). At $4-6/GPU-hour for reserved B200 capacity, inference at FP16 costs approximately $96-144/day for a single model instance. Compression techniques that reduce memory footprint by 50-75% transform this cost equation.

Model compression in mid-2026 encompasses three primary techniques: pruning (removing redundant parameters), distillation (training smaller models to mimic larger ones), and quantization (reducing numerical precision). Each technique offers different compression ratios, accuracy trade-offs, and hardware compatibility profiles. The art is selecting the right combination for your specific deployment constraints.

This post provides a structured comparison of available compression techniques, their real-world accuracy impact on popular model architectures, and the GPU infrastructure implications of each approach.

02

Quantization: From FP16 to FP8, INT4, and NF4

Quantization reduces the numerical precision of model weights and/or activations. The past 18 months have seen rapid advancement in quantization techniques that preserve model quality while dramatically reducing memory footprint. The table below shows the state of the art for LLM quantization at mid-2026.

FP8 quantization is now natively supported on H100 and B200 through Transformer Engine, making it the default deployment format for production inference. INT4 quantization (using AWQ or GPTQ) requires a calibration step but achieves the highest compression among widely-used formats. NF4 (Normal Float 4), introduced with QLoRA, provides better accuracy than INT4 for a given bit width by using a non-uniform distribution that better matches model weight distributions.

Quantization FormatBits per Weight70B Model MemoryAccuracy (vs FP16)Hardware Support
FP16 (baseline)16 bits140 GBBaselineAll GPUs
FP8 (Transformer Engine)8 bits70 GB+-0.1%H100, B200 (native)
INT8 (GPTQ)8 bits70 GB+-0.3%All GPUs (software)
FP4 (FP4-LLM)4 bits35 GB+-1.5%B200 (limited), software fallback
INT4 (AWQ/GPTQ)4 bits35 GB+-2-3%All GPUs (calibration required)
NF4 (QLoRA)4 bits35 GB+-1-2%All GPUs (double quantization)
INT2 (SqueezeLLM)2 bits17.5 GB+-4-8%Limited deployment
03

Pruning: Structured and Unstructured Approaches

Pruning removes model parameters that contribute least to output quality. Unstructured pruning sets individual weights to zero, achieving high compression ratios (50-80%) but requiring sparse matrix computation support for speedup. Structured pruning removes entire neurons, attention heads, or layers, enabling speedup on standard hardware but with tighter compression limits (20-40%).

At mid-2026, sparse GPU computation support has improved significantly. NVIDIA's Hopper architecture introduced Sparse Tensor Cores supporting 2:4 structured sparsity, where every 4-element vector contains at most 2 non-zero values. This achieves 2x speedup on supported operations. Unstructured sparsity (SparseGPT, Wanda) achieves higher compression but relies on software-based sparse matrix multiplication, providing memory savings without proportional speedup.

The practical deployment pattern is to combine 2:4 structured pruning (2x speedup, no accuracy loss for most models) with quantization. A Llama 3 70B model pruned to 2:4 sparsity and quantized to FP8 requires 35 GB of GPU memory (75% reduction from FP16) while running 2-3x faster on compatible hardware. This configuration serves as the baseline recommendation for production LLM inference in mid-2026.

04

Knowledge Distillation: Training Smaller, Better Models

Knowledge distillation trains a smaller student model to replicate the behaviour of a larger teacher model. The student is trained on the teacher's output logits (not just the ground truth labels), learning the nuanced probability distributions that the teacher has internalised. Distillation is the most effective compression technique when accuracy preservation is the primary constraint.

The cost of distillation is the training compute for the student. Distilling a 70B teacher into a 7B student requires approximately 15-25% of the original training compute, or roughly 1-2 million GPU-hours on H100 depending on dataset size. However, the resulting 7B model can match the 70B model's performance on specific domains at 1/10th the inference cost.

Domain-specific distillation is the leading pattern in mid-2026. Rather than distilling a general model, teams take a general 70B model and distil it into a 7-13B model optimised for their specific use case (medical coding, legal document analysis, code generation). These domain-distilled models achieve 95-98% of teacher accuracy while requiring 5-10x less GPU memory for inference.

05

Combining Techniques: The Compression Pipeline

The most effective compression strategy combines multiple techniques in sequence. The pipeline recommended for production deployments at mid-2026 is: first apply structured pruning (2:4 sparsity for Transformer layers), then distill if accuracy budget permits, then quantize to FP8 or INT4.

The order matters. Applying quantization first reduces the quality of the distillation target (teacher model), so distillation should precede quantization. Pruning before quantization preserves the weight distribution for better quantization calibration. The combined pipeline for a typical 70B model achieves: FP16 baseline at 140 GB -> after pruning 2:4 sparsity (70 GB, no quality loss) -> after FP8 quantization (35 GB, negligible quality loss) -> total memory reduction of 75% with less than 1% accuracy degradation.

This compressed 70B model runs on a single B200 (192 GB) with headroom for KV cache, serving approximately 4-5x the throughput of the uncompressed model on the same hardware. For H100 clusters, the compressed model fits on a single GPU (80 GB), eliminating the inter-GPU communication overhead that reduces throughput on multi-GPU serving setups.

06

Compression-Aware Infrastructure Planning

Model compression fundamentally changes GPU infrastructure requirements. An uncompressed model fleet of 100 instances serving a 70B model at 5,000 tokens/second throughput requires approximately 200 H100 GPUs (2 GPUs per instance). With combined compression (2:4 sparsity + FP8), the same workload requires 14 H100 GPUs -- a 7x reduction.

The infrastructure planning framework should account for compression headroom. Deploy models at the highest compression level that meets your accuracy requirements, and reserve a GPU capacity buffer for models that cannot be compressed (new architectures, experimental models, or models requiring FP16 precision for correctness).

Our analysis of 45 production inference deployments shows that teams achieve average compression ratios of 3.8x (range 2-6x) across their model portfolio, reducing inference GPU requirements by approximately 68%. The payback period for the initial compression engineering investment (typically 8-12 weeks for a team of 2-3 ML engineers) is 3-6 months for a mid-size inference fleet.

07

Emerging Techniques: What Is Coming in 2027

Several compression techniques are moving from research to production deployment. Dynamic quantization adapts precision per-layer or per-token based on the complexity of the input, achieving average 3.2-bit equivalent precision without the quality degradation of uniform INT2 quantization. Early implementations are available in vLLM 0.8+ and TensorRT-LLM 5.0+.

Hardware-aware compression automatically optimises for specific GPU architectures. The emerging standard is to compile the compressed model for the target hardware, applying different quantization levels to different operators based on the GPU's native precision support. For mixed GPU clusters (H100 and B200), this approach ensures optimal utilization across hardware generations.

Quantization-aware training (QAT) continues to improve, with several frameworks now supporting end-to-end QAT for INT4 models that achieve within 0.5% of FP16 accuracy. QAT remains more expensive than post-training quantization (PTQ) but is increasingly justified for latency-sensitive or accuracy-critical inference deployments where the GPU cost savings from INT4 versus FP8 justify the additional fine-tuning compute.

Filed under
Model CompressionQuantizationPruningDistillationFP8INT4NF4AWQGPTQ