QUANTIZATION METHOD PROFILES AND MEMORY REDUCTION
Each quantization method makes different GPU hardware assumptions. AWQ uses activation-aware weight quantization that identifies 1% of weight channels as salient and scales them before quantization, achieving W4A16 with minimal quality loss. GPTQ applies approximate second-order (Hessian-based) quantization layer by layer, optimizing the weight grid for each layer's input distribution. GGUF (llama.cpp format) uses k-quant quantization with block sizes of 32 and variable bitwidth per group, optimized for CPU offloading but increasingly run on GPU via llama.cpp's CUDA backend. Native FP8 (transformer engine) is the newest player, leveraging H100 and B200's hardware tensor core support for FP8 matmul.
The VRAM reduction is significant across methods: Llama 3.1 70B at FP16 requires 140 GB. At W4A16 (4-bit weights, 16-bit activations), AWQ and GPTQ reduce this to 38 GB + 8 GB for KV cache at 8K context = 46 GB, fitting on a single H100. GGUF Q4_K_M uses 4.5-bit effective quantization, producing 42 GB on disk and 44 GB in GPU memory with KV cache. FP8 (W8A8) reduces to 70 GB + 8 GB = 78 GB, barely fitting on one H100. The difference between 46 GB and 78 GB determines whether a model fits on one GPU or needs two, with massive cost implications: a single H100 at $2.50/hr versus 2x H100 at $5.00/hr.
| Method | Precision | Llama 70B VRAM | Llama 8B VRAM | Quality Loss (MMLU) | GPU Compatibility |
|---|---|---|---|---|---|
| FP16 (Baseline) | W16A16 | 140 GB + KV | 16 GB + KV | 0% | All GPUs |
| FP8 (Native) | W8A8 | 70 GB + KV | 8 GB + KV | -0.1% | H100, B200 only |
| AWQ | W4A16 | 38 GB + KV | 5 GB + KV | -0.3 to -0.5% | All GPUs (CUDA) |
| GPTQ | W4A16 | 38 GB + KV | 5 GB + KV | -0.4 to -0.8% | All GPUs (CUDA) |
| GGUF Q4_K_M | W4.5A16 | 42 GB + KV | 6 GB + KV | -0.5 to -1.0% | All GPUs + CPU |
| GGUF Q2_K | W2.4A16 | 24 GB + KV | 3.5 GB + KV | -2.5 to -4.0% | All GPUs + CPU |
THROUGHPUT AND LATENCY COMPARISON
Throughput varies significantly by quantization method due to dequantization overhead and kernel support. On a single H100 SXM with Llama 3.1 70B at batch size 1, AWQ achieves 42 tok/s, GPTQ achieves 38 tok/s, GGUF Q4_K_M achieves 31 tok/s, and FP8 achieves 58 tok/s. FP8's 38% advantage over AWQ comes from hardware-native FP8 tensor core execution requiring no dequantization step. At batch size 16, the gap narrows: AWQ 380 tok/s, FP8 420 tok/s (10% advantage), because the dequantization overhead is amortized across larger batches. GGUF's CUDA backend shows the worst performance at all batch sizes due to its block-wise dequantization kernels that are less optimized than the group-wise AWQ kernels.
For Llama 3.1 8B on a single L40S, the differences are smaller because memory is not the bottleneck. AWQ achieves 490 tok/s, GPTQ 470 tok/s, GGUF Q4_K_M 410 tok/s, and FP16 420 tok/s at batch size 32. The smaller model fits in VRAM even at FP16 (16 GB + 2 GB KV), so quantization only helps with throughput by reducing memory bandwidth pressure. On A100 80 GB PCIe, AWQ 70B at batch size 4 achieves 210 tok/s versus FP8 at 280 tok/s on H100, showing that FP8-native hardware acceleration gives H100 a 33% throughput advantage beyond raw memory bandwidth differences.
| Config | BS=1 tok/s | BS=16 tok/s | Max BS | P99 Latency BS=1 | Cost/1M tok |
|---|---|---|---|---|---|
| FP16 70B, 1x H100 | N/A (OOM) | N/A | 0 | - | - |
| FP8 70B, 1x H100 | 58 | 420 | 8 | 1,210 ms | $0.83 |
| AWQ 70B, 1x H100 | 42 | 380 | 12 | 1,540 ms | $0.81 |
| GPTQ 70B, 1x H100 | 38 | 350 | 10 | 1,680 ms | $0.89 |
| GGUF Q4_K_M 70B, 1x H100 | 31 | 280 | 14 | 2,100 ms | $1.02 |
| AWQ 70B, 2x H100 | 88 | 720 | 24 | 780 ms | $0.87 |
QUALITY VS COMPRESSION TRADEOFF ANALYSIS
The quality loss from quantization is task-dependent and asymmetric. On MMLU-Pro (knowledge + reasoning), AWQ W4A16 Llama 3.1 70B loses 0.3% versus FP16 baseline, GPTQ loses 0.6%, and GGUF Q4_K_M loses 0.9%. On GSM8K (math reasoning), the differences widen: AWQ -0.8%, GPTQ -1.4%, GGUF Q4_K_M -2.1%. On HumanEval (code generation), AWQ loses 0.5%, GPTQ loses 1.1%, GGUF loses 1.8%. FP8 shows the smallest degradation across all benchmarks at 0.1-0.3%, making it the preferred quantization for production deployments on H100 and B200 where maintainin g output quality is critical.
The cost-quality Pareto frontier shifts by model size. For 8B models, AWQ is almost always optimal: 5 GB weight memory (versus 16 GB FP16), 0.2% quality loss, and 490 tok/s throughput on L40S. For 70B models, FP8 offers the best quality at 70 GB (fits single H100) versus AWQ at 38 GB with 0.3% more loss. For 405B models, AWQ at 205 GB (fits on 3x H100 at $7.50/hr) versus FP8 at 405 GB (5x H100 at $12.50/hr) makes AWQ the only economically viable option. The decision rule: use FP8 when the model fits on the target GPU count at that precision; use AWQ when it doesn't; use GPTQ when deploying on A100 systems where AWQ kernel support is limited; use GGUF only when CPU offloading is required.
GPU-SPECIFIC QUANTIZATION PERFORMANCE
GPUs handle quantization differently due to architectural variations. H100 SXM with FP8 tensor cores achieves 2x throughput for FP8 inference versus W4A16, because 8-bit matrix operations use the H100 Transformer Engine's accelerated path. A100 has no FP8 support, so AWQ W4A16 is the most performant quantization: 210 tok/s for 70B at BS=4 on A100 80GB PCIe, versus FP16 which requires 2 GPUs at 2x $1.75/hr = $3.50/hr. L40S with 4th-gen tensor cores supports FP8 but at lower throughput than H100: 35 tok/s for 70B at BS=1 with FP8, versus 48 tok/s with AWQ. On L40S, AWQ outperforms FP8 because the dequantization overhead is less than the bandwidth cost of loading 70 GB instead of 38 GB.
The GGUF ecosystem is uniquely suited for hybrid GPU/CPU execution. On a system with a single L40S (48 GB) and 256 GB CPU RAM, GGUF Q4_K_M Llama 3.1 70B runs with 42 GB on GPU and the remaining 2 GB + KV cache 4 GB = 6 GB headroom. When batch size exceeds GPU memory (above BS=6), GGUF seamlessly offloads KV cache to CPU RAM via the llama.cpp CUDA backend. AWQ and GPTQ do not support automatic offloading, crashing on OOM. For GPU-limited deployments with bursty traffic, GGUF's graceful degradation through CPU offloading provides reliability advantages. However, latency during offloading increases by 150-300 ms per request, making it unsuitable for interactive applications below 2-second latency targets.
