All essays
BenchmarkCOMPARISONFEB 2026

ONNX Runtime vs TensorRT vs OpenVINO: Cross-Platform GPU Inference Comparison

Production benchmark comparing ONNX Runtime, TensorRT, and OpenVINO for cross-platform GPU inference. NVIDIA vs Intel vs AMD GPU support, FP8/INT4 quantization, throughput benchmarks on H100, A100, Arc, and MI350X.

01

GPU ARCHITECTURE SUPPORT AND OPTIMIZATION DEPTH

The three inference optimization frameworks diverge in their GPU vendor support. NVIDIA TensorRT is NVIDIA-only, supporting CUDA GPUs from Maxwell (sm_52) through Blackwell (sm_100) with the deepest optimization for H100 and B200. TensorRT 10.7+ includes the TensorRT-LLM plugin for LLM-specific optimizations: FP8 GEMM with per-tensor scaling, in-flight batching, paged KV cache, and speculative decoding. ONNX Runtime (ORT) offers the widest GPU support through its execution providers: CUDA (NVIDIA), ROCm (AMD), DML (DirectML for all Windows GPUs including Intel Arc and NVIDIA), OpenVINO (Intel), and CoreML (Apple). ORT's extensible provider architecture enables a single ONNX model to run across GPU vendors without modification. OpenVINO is Intel-centric: optimized for Intel integrated GPUs, Intel Arc discrete GPUs, and Intel MAX GPUs, with limited NVIDIA CUDA support via the default plugin.

The optimization depth varies by framework. TensorRT applies the most aggressive optimizations: layer fusion across 150+ operator patterns, automatic FP8/INT8 calibration with per-channel quantization, kernel autotuning with 50-200 GEMM variants per operation, and memory pool optimization that reduces peak usage by 20-35%. ORT's optimizations are framework-level: graph transformation (constant folding, node elimination), operator fusion (GELU fusion, LayerNorm fusion), and execution provider delegation. ORT does not perform quantization calibration automatically; that requires ONNX Runtime's quantization tool or a separate calibration pass. OpenVINO's optimization strengths are in model compression: automatic INT8 and INT4 weight compression with 1% quality loss using its NNCF (Neural Network Compression Framework), and a unique FP16-to-INT4 progressive compression pipeline for memory-constrained deployments.

FeatureTensorRT 10.7ONNX Runtime 1.20OpenVINO 2025.4
GitHub Stars~9,500 (repo)~14,000~6,800
NVIDIA GPU SupportFull (sm_52 to sm_100)CUDA EP (sm_60+)Limited (default plugin)
AMD GPU SupportNoneROCm EP (MI100+)None
Intel GPU SupportNoneDML EP (Arc/Ultra)Full (Arc, MAX, iGPU)
FP8 QuantizationNative (H100+)Via QDQ nodesVia NNCF compression
INT4 QuantizationVia AWQ/GPTQ pluginVia ONNX quantizationVia NNCF progressive
02

THROUGHPUT AND LATENCY BENCHMARKS ACROSS GPUS

Benchmarking Llama 3.1 8B on H100 SXM 80 GB at batch size 1 with FP16 shows TensorRT-LLM achieving 420 tok/s, ORT with CUDA EP achieving 340 tok/s (19% slower), and OpenVINO with the GPU plugin achieving 210 tok/s (50% slower). The gap narrows at batch size 32: TensorRT-LLM at 8,200 tok/s, ORT at 7,100 tok/s (14% slower), and OpenVINO at 4,800 tok/s (41% slower). TensorRT's advantage comes from its custom CUDA kernel implementations (FlashAttention-2 integrated into the TRTLLM plugin, fused MoE gate computation, and paged KV cache with optimized memory management). ORT's CUDA EP relies on PyTorch's underlying CUDA kernels, losing the custom optimization opportunity but gaining ease of deployment.

On AMD MI350X, ORT with ROCm EP achieves 290 tok/s for Llama 8B at BS=1 versus TensorRT (not available on AMD) and OpenVINO (not available on AMD). On Intel Arc A770 (16 GB), OpenVINO achieves 82 tok/s for Llama 8B at BS=1, ORT with DML EP achieves 65 tok/s, and TensorRT is not available. The INT4 quantization levels the playing field: with OpenVINO NNCF INT4 compression on Arc A770, Llama 8B throughput improves to 145 tok/s, a 1.8x improvement over FP16, making Intel GPUs viable for low-cost inference. For cross-platform GPU deployments serving diverse hardware in a single cluster, ORT is the pragmatic choice because it supports all vendors with a single ONNX model file, though the performance penalty versus TensorRT on NVIDIA hardware is 14-19%.

Model + GPUTensorRT-LLMONNX RuntimeOpenVINOBest Performer
Llama 8B, H100 BS=1420 tok/s340 tok/s210 tok/sTensorRT-LLM +19%
Llama 8B, H100 BS=328,200 tok/s7,100 tok/s4,800 tok/sTensorRT-LLM +14%
Llama 8B, A100 BS=1280 tok/s240 tok/s160 tok/sTensorRT-LLM +14%
Llama 8B, MI350X BS=1N/A290 tok/sN/AORT (only option)
Llama 8B, Arc A770 BS=1N/A65 tok/s82 tok/sOpenVINO +21%
ResNet-50, H100 BS=25648,000 img/s41,000 img/s32,000 img/sTensorRT +15%
03

QUANTIZATION PIPELINES AND MODEL CONVERSION WORKFLOW

The model conversion workflow differs substantially between frameworks. TensorRT requires converting PyTorch or ONNX models into TRT engines via `trtexec` or the TensorRT Python API. The conversion pipeline: export model to ONNX via `torch.onnx.export()`, run `trtexec --onnx=model.onnx --fp8 --saveEngine=model.plan`, then load the plan file in the serving runtime. The conversion takes 5-60 minutes for a 70B model, depending on the number of autotuning iterations. TensorRT-LLM simplifies this with `trtllm-build` that directly accepts Hugging Face model paths and produces TRTLLM engines in a single command with built-in INT4 AWQ weight quantization: `trtllm-build --model_dir ./llama-70b --weight_only_precision int4_awq --output_dir ./trtllm_engine`.

ONNX Runtime's conversion is simpler but less optimized: export to ONNX, optionally quantize via ONNX Runtime's `quantize_dynamic` or `quantize_static` API, then run the ONNX model through ORT's execution providers. Unlike TensorRT's aggressive per-channel quantization calibration, ORT's quantization is per-tensor by default, resulting in 1-3% higher quality loss at INT8. OpenVINO's conversion uses `optimum-intel` package: `optimum-cli export openvino --model meta-llama/Llama-3.1-8B --weight-format int4 ./openvino_model`. The OpenVINO INT4 compression with NNCF's mixed-precision algorithm achieves quality within 0.5% of FP16 while reducing model size by 4x. For multi-vendor GPU clusters, the recommended deployment pipeline converts the model to each framework separately and selects the appropriate engine at serving time based on the node's GPU vendor.

04

PRODUCTION DEPLOYMENT PATTERNS FOR MULTI-VENDOR CLUSTERS

For GPU clusters with mixed vendors (NVIDIA, AMD, Intel), the production pattern uses a model-gateway architecture. The gateway (typically Envoy or a custom gRPC proxy) inspects the target GPU node's vendor label and routes the request to the appropriate inference runtime: TensorRT-LLM for NVIDIA nodes, ORT with ROCm EP for AMD nodes, OpenVINO for Intel nodes. Each node preloads the model in its native format and exposes a unified gRPC or HTTP API. The gateway maintains a vendor-to-runtime mapping registered at node startup via Kubernetes annotations. This architecture supports up to 15% throughput variation across nodes while maintaining a single serving endpoint for clients.

On ClusterBid, the most cost-effective multi-vendor configuration for inference uses 80% H100/TensorRT-LLM for high-throughput workloads and 20% L40S/L4 (NVIDIA) or MI350X (AMD with ORT) for burst capacity and latency-tolerant batch workloads. The burst nodes cost $1.10-1.80/hr versus $2.50/hr for H100, providing 30-56% cost savings for non-latency-critical traffic. The TensorRT-LLM engine compilation must be included in the CI/CD pipeline as a build step with GPU CI runners, adding 15-30 minutes to the deployment pipeline for engine conversion. For teams deploying on single-vendor GPU clusters, TensorRT on NVIDIA or OpenVINO on Intel provides the highest performance; for multi-vendor flexibility, ONNX Runtime is the practical default despite the 14-19% performance penalty on NVIDIA hardware.

Filed under
ONNX Runtime GPUTensorRT OptimizationOpenVINO GPUCross-Platform InferenceGPU Inference BenchmarksFP8 TensorRTMulti-Vendor GPU Inference