QUANTIZATION METHODS FOR GPU DEPLOYMENT
Quantization reduces model precision from FP32 to lower bit-widths, delivering the largest memory and latency gains. NVIDIA H100 Tensor Cores support FP8 natively at 2x throughput versus FP16, while INT4 via AWQ or GPTQ achieves 4x memory reduction. Llama 3 70B in FP16 requires 140 GB of GPU memory, fitting only across two H100 GPUs. INT4 quantization compresses to 35 GB, fitting on a single H100 and achieving 2.8x higher throughput.
Weight-only quantization (W4A16) stores weights in INT4 but computes in FP16, preserving activation precision. SmoothQuant and AWQ calibrate quantization scales using 128-512 calibration samples, reducing perplexity loss below 0.2 points on Llama 3 models. Per-group quantization with group size 128 achieves 2.5-3.0x memory reduction with less than 0.3 percent accuracy loss versus full FP16.
| Precision | Memory/Model (70B) | Speedup vs FP32 | Accuracy Loss | GPU Support |
|---|---|---|---|---|
| FP32 | 280 GB | 1.0x | Baseline | All GPUs |
| FP16/BF16 | 140 GB | 2.0x | <0.1% | V100+ Tensor Cores |
| FP8 | 70 GB | 4.0x | 0.2-0.5% | H100/H200/B200 |
| INT4 (GPTQ) | 35 GB | 6.0-8.0x | 0.5-1.5% | All >= A100 |
| INT4 (AWQ) | 35 GB | 6.5-8.5x | 0.3-1.0% | All >= A100 |
KNOWLEDGE DISTILLATION FOR DEPLOYMENT
Knowledge distillation trains a smaller student model to replicate a larger teacher model behavior. Distilling Llama 3 70B into a 7B student achieves 10x inference speedup while retaining 92-95 percent of task-specific accuracy. The process requires generating 500K-2M synthetic training pairs consuming approximately 8,000-32,000 H100-hours at a cost of $28,000-$112,000. ROI is positive within 3-4 months for deployments serving over 100 million tokens daily.
Domain-specific distillation outperforms general distillation by 3-8 percentage points. A code generation model distilled on 1 million code-completion pairs retains 96 percent of CodeLlama 34B accuracy in a 7B model versus 89 percent with general distillation. Temperature scaling during distillation with T=2.0 to T=4.0 produces softer probability distributions that improve student learning.
| Compression Method | Model Size Reduction | Speedup | Accuracy Retention | Training Cost |
|---|---|---|---|---|
| INT4 Quantization | 4x | 6-8x | 98.5-99.7% | $0 (post-hoc) |
| Knowledge Distillation | 10x | 10x | 92-95% | $28K-$112K |
| Structured Pruning | 2-3x | 2-3x | 95-98% | $5K-$20K |
| Unstructured Pruning | 5-10x | 1-2x* | 90-97% | $2K-$10K |
| Combined (Distill + Quant) | 40x | 12-16x | 88-93% | $30K-$120K |
PRUNING STRATEGIES FOR GPU INFERENCE
Structured pruning removes entire attention heads, layers, or channels, producing models that run efficiently on GPU hardware without sparse matrix support. Removing 25 percent of attention heads from Llama 3 models reduces inference latency by 18-22 percent with 1-2 percent accuracy loss. Layer pruning removing 4 of 32 layers reduces throughput by 12 percent but achieves 25 percent latency improvement.
GPU hardware support for sparsity varies significantly. NVIDIA Ampere and Hopper architectures support 2:4 structured sparsity yielding 2x theoretical speedup. In practice real-world speedup is 1.3-1.6x due to memory bandwidth constraints. A100 achieves 1.5x speedup with 2:4 sparsity while H100 achieves 1.6x. Sparsity-aware kernels from NVIDIA cuSPARSELt are required for inference.
PRODUCTION COMPRESSION PIPELINE
A production compression pipeline follows a four-stage workflow. Stage one profiles the model to identify compression sensitivity. Stage two applies quantization with 128-512 calibration samples. Stage three optionally applies structured pruning or distillation if accuracy targets are not met. Stage four validates the compressed model against a holdout set of 10,000 samples.
Continuous monitoring of compressed model quality in production is essential. Drift detection comparing output distributions of compressed and original models should trigger re-calibration when KL divergence exceeds 0.05. A deployment serving 5 million tokens daily should re-calibrate every 7-14 days. Companies using automated compression pipelines report 94-97 percent deployment success rates.
