All essays
BenchmarkCOMPARISONFEB 2026

ONNX Runtime vs TensorRT vs OpenVINO: GPU Inference Engine Benchmark for Production AI Serving

Head-to-head benchmark of ONNX Runtime, NVIDIA TensorRT, and Intel OpenVINO inference engines on H200 and B200 GPUs, covering latency, throughput, model coverage, and deployment cost per inference.

01

The Inference Engine Landscape in 2026

Three inference engines dominate GPU-accelerated model serving in 2026. ONNX Runtime (Microsoft, open source, v1.22 as of June 2026) provides a cross-platform execution layer supporting CUDA, ROCm, OpenVINO, DirectML, and WebGPU execution providers. NVIDIA TensorRT (v10.8) is the proprietary engine optimized exclusively for NVIDIA GPUs, supporting the full range of quantization modes from FP16 through FP8, INT8, INT4, and FP4 on Blackwell architectures. Intel OpenVINO (v2026.2) has expanded beyond Intel hardware to support CUDA and ROCm backends while maintaining its lead on Intel GPU and CPU inference.

The choice between engines depends on GPU type, model architecture, quantization requirements, and deployment platform. TensorRT is the default for NVIDIA-only deployments. ONNX Runtime is preferred for multi-platform or multi-vendor GPU fleets. OpenVINO is the primary choice for Intel GPU (Flex, Max) and hybrid CPU/GPU deployments. Each engine offers different tradeoffs in latency, throughput, model op coverage, kernel optimization depth, and integration complexity with production serving frameworks like Triton or KServe.

02

ONNX Runtime Architecture

ONNX Runtime converts a model from any framework (PyTorch via torch.onnx.export, TensorFlow via tf2onnx, JAX via jax2tf) into the ONNX intermediate representation, then applies graph optimizations and delegates computation to the optimal execution provider for the target hardware. The CUDA execution provider (v1.22) supports Tensor Core operations, fused attention, FlashAttention-2 integration, and FP8 quantization via NVIDIA Transformer Engine. The ROCm execution provider enables the same ONNX model to run on AMD MI300X/325X GPUs without code changes.

The key advantage of ONNX Runtime is model portability. An ONNX model exported once with dynamic axes can be deployed to NVIDIA GPUs via CUDA EP, AMD GPUs via ROCm EP, Intel GPUs via OpenVINO EP, or CPUs via the default MLAS provider. The tradeoff is that ONNX Runtime's CUDA kernel coverage is less exhaustive than TensorRT's, resulting in 10-30% lower peak throughput for models with non-standard operations like flash attention v3 kernels, grouped query attention with large group counts, or custom activation functions. For standard transformer architectures (LLaMA, GPT, BERT), ONNX Runtime CUDA EP achieves 85-95% of TensorRT throughput.

03

TensorRT Optimization Pipeline

TensorRT's optimization pipeline applies five passes: graph optimization (node fusion, constant folding, dead-code elimination), precision calibration (FP8/INT8/INT4 with per-tensor or per-channel scaling factors), kernel autotuning (selecting the optimal CUDA kernel for each operation given the target GPU's SM count, memory bandwidth, and Tensor Core generation), memory planning (reuse of scratch buffers via memory pool optimization), and serialization to a TensorRT engine file (plan file) for deployment. The engine build is model-specific and GPU-specific: a TensorRT engine built for B200 FP4 cannot run on H200 FP8 and must be rebuilt.

The B200's Transformer Engine v2 integration in TensorRT 10.8 provides native FP4 matrix operations, FP4 KV-cache quantization, and second-order tensor core scheduling that overlaps memory transfers with computation. TensorRT 10.8 delivers an average 1.8x throughput improvement over TensorRT 10.6 on B200 for transformer inference. The primary limitation is that TensorRT only supports NVIDIA GPUs. A deployment that needs to serve the same model across NVIDIA, AMD, and Intel GPUs must maintain separate model artifacts and deployment pipelines for each engine, increasing operational complexity.

04

OpenVINO Architecture

Intel OpenVINO 2026.2 converts models to its Intermediate Representation (IR) format via the Model Optimizer or direct ONNX import. The inference is executed through plugins that map to target hardware. The CUDA plugin (added in 2024.1) provides an alternative execution path, though it does not support the full OpenVINO opset and lacks optimization for NVIDIA-specific features like FP4 or INT4 AWQ. OpenVINO's strength is on Intel hardware: Intel Flex 170 GPUs, Arc A770, and the upcoming Intel Falcon Shores GPU, where it achieves 1.2-1.4x the throughput of ONNX Runtime on the same hardware.

For multi-platform inference serving, OpenVINO serves as the bridge between Intel and non-Intel hardware. A model deployed via OpenVINO can run unchanged on Intel CPUs (using the CPU plugin with VNNI instructions for INT8 inference), Intel GPUs (using the GPU plugin with XMX acceleration), and NVIDIA GPUs (using the CUDA plugin). The CUDA plugin performance on H200 for transformer models is approximately 60-75% of TensorRT throughput and 70-85% of ONNX Runtime CUDA EP throughput. OpenVINO is the best choice for deployments that run primarily on Intel hardware with occasional NVIDIA GPU burst capacity.

05

Latency Benchmarks on H200 and B200

The following latency benchmarks use LLaMA-3 8B at FP16 with input length 2048 and output length 256, measured on a single GPU with batch size 1. TensorRT delivers the lowest median time-to-first-token on both H200 and B200, with ONNX Runtime CUDA EP within 12-18% and OpenVINO CUDA plugin at 35-45% higher TTFT. The B200's advantage over H200 is most pronounced in TTFT due to the higher HBM3e bandwidth and faster tensor core architecture.

Time-per-output-token (TPOT) shows a similar pattern. TensorRT is the leader on both platforms. ONNX Runtime trails by 15-22% on TPOT depending on whether the model uses grouped query attention (which TensorRT fuses more aggressively). OpenVINO CUDA plugin on B200 delivers approximately 85% of ONNX Runtime TPOT and 65% of TensorRT. The gap between TensorRT and OpenVINO is largest on models that use flash attention v3 kernels, which OpenVINO's CUDA plugin does not support natively.

EngineGPUTTFT (ms)TPOT (ms)Tokens/sRelative to TensorRT
TensorRT 10.8H200428.21221.00x (baseline)
ONNX Runtime 1.22 CUDAH200489.81020.84x
OpenVINO 2026.2 CUDAH2005812.1830.68x
TensorRT 10.8B200285.41851.52x
ONNX Runtime 1.22 CUDAB200336.61521.25x
OpenVINO 2026.2 CUDAB200418.51180.97x
06

Throughput Benchmarks at Production Batch Sizes

At production batch sizes (64 concurrent requests, input 2048, output 512), TensorRT maintains its lead on both H200 and B200. The benchmark uses LLaMA-3 8B at FP8 on H200 and FP4 on B200 (where supported). TensorRT processes 3,480 tokens/s on H200 FP8 and 5,810 tokens/s on B200 FP4. ONNX Runtime achieves 2,940 tokens/s on H200 and 4,720 tokens/s on B200. OpenVINO delivers 2,100 tokens/s on H200 FP8 and 3,560 tokens/s on B200 (using FP8 since OpenVINO does not yet support FP4).

The cost-per-million-tokens at ClusterBid mid-2026 spot rates tells the economic story. On H200 at $3.07/hr, TensorRT costs $0.245 per million tokens, ONNX Runtime costs $0.290, and OpenVINO costs $0.406. On B200 at $5.45/hr, TensorRT FP4 costs $0.260 per million tokens, ONNX Runtime FP8 costs $0.321, and OpenVINO costs $0.425. TensorRT on B200 FP4 is the most cost-effective configuration at approximately $0.26 per million tokens. For multi-platform deployments that sacrifice TensorRT's peak throughput for portability, ONNX Runtime on B200 is the next best option at $0.32 per million tokens.

EngineGPU + QuantThroughput (tok/s)Cost/hr (spot)Cost/M tokModel Coverage
TensorRT 10.8H200 FP83,480$3.07$0.245Excellent (NVIDIA only)
ONNX Runtime 1.22H200 FP82,940$3.07$0.290Excellent (all platforms)
OpenVINO 2026.2H200 FP82,100$3.07$0.406Good (Intel + NVIDIA)
TensorRT 10.8B200 FP45,810$5.45$0.260Excellent (NVIDIA only)
ONNX Runtime 1.22B200 FP84,720$5.45$0.321Excellent (all platforms)
OpenVINO 2026.2B200 FP83,560$5.45$0.425Good (Intel + NVIDIA)
07

Model Coverage and Compatibility

TensorRT 10.8 supports the widest range of operations on NVIDIA GPUs through its TensorRT-LLM integration, including transformer-specific ops (FlashAttention-2 and 3, GQA fusion, MoE routing kernels, FP4 matrix multiply, multi-block sparse attention). The coverage for standard transformer models (LLaMA, Mistral, GPT, BERT, T5, Whisper, Stable Diffusion, DiT) is comprehensive. The main incompatibility is custom operations not covered by TensorRT's plugin registry, which require custom CUDA plugin development.

ONNX Runtime supports the broadest range of hardware platforms. Any model that exports to ONNX (which covers 90%+ of PyTorch, TensorFlow, and JAX models) can run on any execution provider. The coverage limitation is at the execution provider level: not all ONNX ops are supported by every EP. The CUDA EP supports 95%+ of common ONNX ops. The ROCm EP supports approximately 85% of CUDA EP ops. OpenVINO EP covers approximately 90% of CUDA EP ops. Models that use very new operations (like FP4 matmul) will fall back to less optimized paths in ONNX Runtime or OpenVINO until those engines add native support.

08

Inference Engine Selection Guide

Use TensorRT when your entire GPU fleet is NVIDIA (H100, H200, B200, or B300), you need maximum throughput per dollar, and your model ops are covered by TensorRT's plugin registry. TensorRT delivers 15-30% more throughput than ONNX Runtime on NVIDIA GPUs. The single-platform limitation is acceptable for teams that standardize on NVIDIA hardware, which describes approximately 85% of production AI deployments in 2026.

Use ONNX Runtime when your deployment targets multiple GPU vendors (NVIDIA + AMD, or NVIDIA + Intel), you need to port models between platforms without re-exporting, or you are deploying to a heterogeneous fleet with different GPU architectures. The 15-30% throughput penalty versus TensorRT on NVIDIA GPUs is the cost of platform flexibility. ONNX Runtime's ROCm EP makes it the only practical choice for AMD GPU deployments.

Use OpenVINO when Intel GPUs (Flex, Arc, Falcon Shores) form the primary deployment target, or when deploying to hybrid CPU/GPU inference pipelines with Intel CPUs. OpenVINO's CPU inference is best-in-class for INT8 quantized models using Intel VNNI and AMX instructions, often matching GPU throughput for small models (under 3B parameters). For pure NVIDIA GPU deployments, OpenVINO is not recommended. ClusterBid supports all three inference engines with pre-configured deployment templates on H200 and B200 clusters.

Filed under
ONNX RuntimeTensorRTOpenVINOinference benchmarkGPU inferencemodel servinginference engineH200 B200 inference