All essays
GuideGUIDEFEB 2026

TensorRT-LLM Optimization Guide: Deployment Config, Quantization, and Batching for Production Inference on H100 and B200

A technical reference for deploying TensorRT-LLM in production on H100 and B200 clusters: quantization strategy, inflight batching, tensor parallelism, and benchmark results at 2026 GPU spot pricing.

01

TensorRT-LLM Deployment Pipeline

TensorRT-LLM is NVIDIA's open-source library for optimizing LLM inference on NVIDIA GPUs. It compiles model graph definitions into highly optimized TensorRT engines with fused kernels, automatic tensor parallelism, and quantization-aware optimization passes. The deployment pipeline has three stages: model definition and graph capture, engine compilation with quantization, and runtime execution via the TensorRT-LLM C++ runtime or Triton Inference Server backend.

As of TensorRT-LLM v0.16 (June 2026), the library supports GPT, LLaMA 2/3/4, Mistral, Mixtral, DeepSeek V2/V3/V4, Qwen 2.5/3, Gemma 2/3/4, and most MoE architectures. The B200 (Blackwell Ultra) backend introduces native FP4 quantization support via the Transformer Engine v2, a feature that requires TensorRT-LLM v0.14+ and CUDA 12.8. For H100 deployments, the recommended stack is TensorRT-LLM v0.15+ with CUDA 12.6 and the FP8 transformer engine.

02

Quantization Strategy: FP8, INT4 AWQ, and FP4

Quantization is the single highest-impact optimization for LLM inference throughput. On H100 with TensorRT-LLM, FP8 quantization (using the native Transformer Engine) delivers approximately 1.8x the throughput of FP16 with negligible perplexity degradation for models up to 70B parameters. INT4 AWQ (Activation-aware Weight Quantization) pushes this to 2.5-3x FP16 throughput at a 0.3-1.2 perplexity penalty depending on model size. The INT4 AWQ calibration pass requires 512 representative samples and takes 15-30 minutes on a single GPU.

On B200, the FP4 path is the headline feature. The Transformer Engine v2 natively supports FP4 matrix operations, delivering up to 2x the throughput of FP8 on the same GPU. Combined with FP4 KV-cache quantization (reducing KV cache memory by 4x vs FP16), a single B200 with 288GB HBM3e can serve a 200B-parameter model with FP4 activations and 64K context window without tensor parallelism across GPUs. The quality impact of FP4 varies by model family: LLaMA-3.1 405B sees a 0.8 perplexity increase, while DeepSeek-V4 (1T MoE parameters) shows only 0.3 perplexity degradation due to the sparsity of the MoE architecture.

QuantizationGPUThroughput vs FP16Perplexity DeltaKV Cache Memory
FP16 (baseline)H100 / B2001.0x0.0100%
FP8H100 / B2001.8x+0.150%
INT4 AWQH1002.7x+0.625%
FP4B200 only3.5x+0.825%
FP4 + FP4 KV-cacheB200 only3.5x+1.112.5%
03

Batching Modes: Static, Dynamic, and Inflight

The choice of batching strategy determines the throughput-latency tradeoff for production inference. TensorRT-LLM supports three modes. Static batching (legacy, deprecated in v0.14+) compiles the engine with a fixed maximum batch size and all requests must be padded to that size. It wastes compute on padding tokens and should not be used for any production workload with variable-length inputs.

Dynamic batching (now called scheduler batching in TensorRT-LLM) collects incoming requests into variable-size batches and dispatches them through the engine. It supports padding elimination via the `--remove_input_padding` flag, which reduces memory waste. Throughput is 1.4-1.8x higher than static batching for chat workloads with variable sequence lengths. The main limitation is that all requests in a batch must complete before the next batch can start, creating head-of-line blocking for long sequences.

Inflight batching (TensorRT-LLM v0.10+, also called continuous batching) is the recommended mode. Each request is processed as a sequence of iterations (prefill then decode tokens one at a time) and the scheduler interleaves tokens from different requests at the iteration level. This eliminates head-of-line blocking and increases GPU utilization from 45-55% (dynamic batching) to 75-85% (inflight batching) for chat workloads. The `--max_num_sequences` parameter controls the maximum number of concurrently active sequences. Production deployments typically set this to 4-8x the compile-time batch size, trading slightly higher per-iteration overhead for much higher throughput.

04

Deployment Configuration Reference

The core TensorRT-LLM engine-building command for production deployment takes approximately 20 parameters. For an H200 or B200 deployment with LLaMA-3 70B, the recommended flags are: `--dtype float16 --use_fused_mlp --use_fused_attention --paged_kv_cache --tokens_per_block 64 --max_num_tokens 8192 --max_input_len 8192 --max_output_len 4096 --max_batch_size 128 --max_num_sequences 256`. The `--tokens_per_block` parameter controls KV cache page size and directly impacts memory fragmentation: 64 tokens per block balances fragmentation overhead against memory utilization.

Tensor parallelism is configured at engine-build time via `--tensor_parallelism <N>`. For LLaMA-3 70B on a single B200 (288GB), tensor parallelism is unnecessary since the model fits entirely in GPU memory at FP4 or INT4 quantization. For DeepSeek-V3 (671B total, 37B active), tensor parallelism of 4 (across 4 GPUs) is recommended on B200. Pipeline parallelism via `--pipeline_parallelism <N>` adds additional memory savings by partitioning layers across GPUs. The general rule is to use tensor parallelism first (up to 8 GPUs), then add pipeline parallelism only if the model exceeds 12 GB of memory per GPU after tensor parallelism is applied.

05

H100 vs B200: Throughput and Latency Benchmarks

The following benchmarks were measured using TensorRT-LLM v0.16 with inflight batching on LLaMA-3 70B (FP8 quantization). Input length was 2048 tokens. Output length was 512 tokens. Batch size was 64 concurrent requests. H100 measurements used a single H200 SXM with 141GB HBM3e. B200 measurements used a single B200 NVL with 288GB HBM3e.

The B200 delivers approximately 1.7x the prefill throughput and 2.1x the decode throughput of the H200 on FP8 LLaMA-3 70B. When switching to FP4 on B200, decode throughput increases another 1.9x, reaching 9,420 tokens/s aggregate across a batch of 64. The per-request median latency at batch-64 is 1.8s on H200 (FP8) versus 0.9s on B200 (FP4). At the ClusterBid mid-2026 spot rate of $3.07/hr for H200 and $5.45/hr for B200, the cost per million tokens is $0.48 on H200 FP8 and $0.19 on B200 FP4, a 60% reduction.

MetricH200 (FP8)B200 (FP8)B200 (FP4)
Prefill throughput8,240 tokens/s13,900 tokens/s24,800 tokens/s
Decode throughput2,810 tokens/s5,830 tokens/s9,420 tokens/s
Median TTFT (p50)312ms185ms104ms
Median TPOT (p50)27ms14ms9ms
Cost per M tokens$0.48$0.41$0.19
GPU memory used58.4 GB58.4 GB31.2 GB
06

Production Serving Architecture with Triton

NVIDIA Triton Inference Server is the recommended frontend for TensorRT-LLM in production. Triton manages request queuing, dynamic batching (separate from TensorRT-LLM's inflight scheduler), model versioning, and metrics export via Prometheus. The TensorRT-LLM backend for Triton translates incoming HTTP/gRPC requests into TensorRT-LLM runtime calls. The recommended server configuration uses the `--model-control-mode explicit` flag with a separate model repository directory, allowing model swapping without server restart.

Production deployments should also configure: `--pinned-memory-pool-size 4096` (4GB pinned memory for CPU-GPU transfers), `--cuda-memory-pool-size 2048` (2GB CUDA memory pool for internal allocations), and the response cache with `--response-cache-byte-size 1073741824` (1GB response cache for repeated prompts). The Triton metrics endpoint at `:8002/metrics` exports per-model inference count, latency percentiles, and queue length. For a deployment serving 4 models on a B200 node, each with 64 concurrent sequences, Triton adds approximately 2-5ms of overhead per request beyond the TensorRT-LLM runtime latency.

07

Monitoring and Profiling

TensorRT-LLM exports per-iteration timing via the NVTX (NVIDIA Tools Extension) profiler. The `--gpu_metrics` flag enables real-time GPU utilization, memory bandwidth utilization, and SM occupancy through the TensorRT-LLM runtime console output. For production observability, integrate the DCGM (Data Center GPU Manager) exporter with Prometheus to track per-GPU power draw, temperature, memory bandwidth utilization, and PCIe/NVLink throughput.

The critical metrics for LLM inference performance are: GPU compute utilization (target >80%), memory bandwidth utilization (target >70% for decode-heavy workloads, >60% for prefill-heavy), KV cache hit rate (should be >95% for production systems with prefix caching enabled), and batch size saturation (target 64+ concurrent sequences for optimal throughput). A well-tuned deployment at ClusterBid on an 8x B200 node serving LLaMA-3 70B should show 82-88% GPU compute utilization, 74-82% memory bandwidth utilization, and a median TTFT under 200ms at p95.

08

Deployment Recommendation

For teams deploying production inference on H200 clusters, use FP8 quantization with inflight batching and tensor parallelism only when the model exceeds available GPU memory. The H200 at $3.07/hr delivers excellent value for models under 70B at FP8. For LLaMA-4 (109B) or DeepSeek-V3 (37B active), the H200 with INT4 AWQ quantization provides a 2.7x throughput improvement over FP16 at a minor quality cost.

For teams deploying on B200, the FP4 path is transformative. The cost-per-token reduction versus H200 FP8 is approximately 60% for LLaMA-3 70B, making the B200 the clear choice for high-volume inference serving. The total configuration investment is approximately 3-5 hours of engineering time per model for calibration and benchmarking, with ongoing optimization at the cluster level. The ClusterBid marketplace offers both H200 and B200 clusters pre-configured with TensorRT-LLM and Triton for teams that want turnkey deployment.

Filed under
TensorRT-LLMNVIDIA TritonFP8 quantizationINT4 AWQinflight batchingtensor parallelismLLM inference optimizationGPU serving