TensorRT-LLM Overview
TensorRT-LLM (TensorRT-LLM 0.13) by NVIDIA provides GPU-accelerated inference with features: FP8/FP4/INT4, in-flight batching, multi-node, NIM, ONNX, in-flight batching, plugins... It achieves 93-98% GPU utilization, 10-30ms prefill, best peak performance throughput on H100 GPUs. License: Open-source, NVIDIA license, NIM commercial.
Performance Benchmarks
On H100 80GB with Llama 4 Scout (17B) at FP8: prefill throughput: 21,230 tokens/second; decode throughput: 1291 tokens/second per user with 1678 max batch; TTFT (time to first token): 23ms; inter-token latency: 23ms. GPU utilization: 93-98% GPU utilization.
Feature Comparison
Key features: FP8/FP4/INT4, in-flight batching, multi-node, NIM, ONNX, in-flight batching, plugins... Unique strengths: FP8/FP4/INT4 in-flight batching multi-node. Production features include: health checks, metrics endpoints, graceful shutdown.
Cost-Per-Token Analysis
Cost-per-token on H100 80GB at $2.50/hr: input tokens: $0.000053/1K tokens; output tokens: $0.000143/1K tokens with TensorRT-LLM. At 50% utilization, cost-per-million tokens: $169-$339 for output tokens, depending on batch size and model size. Reserved pricing reduces costs by 30-50%.
Production Deployment
Deploy TensorRT-LLM in production: containerized deployment with Docker + NVIDIA Container Toolkit; Kubernetes with GPU node pools; monitoring with Prometheus + GPU metrics; horizontal scaling with HAProxy + keepalived; and CI/CD integration for model updates. Recommended: 5x H100/B200 GPUs per node with NVLink.
When to Choose TensorRT-LLM
Choose TensorRT-LLM when: FP8/FP4/INT4 in-flight batching are critical for your workloads; Open-source, NVIDIA license, NIM commercial license model fits your budget; and your team has experience with NVIDIA's ecosystem. Consider alternatives when: specific hardware optimization is needed, team familiarity with other frameworks, or license costs are prohibitive for your scale.
