All essays
TechnicalDEEP DIVEFEB 2026

AI Inference at the Edge: Jetson Orin, AGX, and Embedded GPU Deployments

AI inference at the edge with NVIDIA Jetson Orin, AGX, and embedded GPUs. Performance benchmarks, power efficiency, and deployment strategies for edge AI.

01

Edge vs Cloud Inference: When to Run Locally and When to Call the Cloud

The decision to deploy AI inference at the edge versus running in the cloud on GPUs like H100 ($1.15/hr on ClusterBid) is a tradeoff of latency, bandwidth, privacy, and total cost of ownership. Cloud inference provides unlimited model capacity: you can serve Llama 4 400B or any large model from a GPU cluster with 1.8 TB/s NVLink interconnects. Edge inference is constrained by the embedded GPU's capabilities - a Jetson AGX Orin 64GB delivers roughly 275 TOPS of INT8 throughput, sufficient for models up to 7-13B parameters with aggressive quantization, but not for 70B+ models.

The latency argument for edge AI is decisive for real-time applications. Cloud inference adds 30-100ms of network round-trip time to every request, regardless of the GPU inference speed. For autonomous systems - robotics, drones, autonomous vehicles, industrial automation - 100ms of additional latency is unacceptable for control loops that must respond in milliseconds. These applications require edge inference where the model runs on the same device as the sensors and actuators, with inference latency measured in microseconds to low milliseconds.

The bandwidth and privacy arguments favor edge inference for applications that process sensitive data or generate high-volume outputs. A security camera system analyzing 30 frames per second from 100 cameras generates 162GB of video data per day per camera if pushed to the cloud for inference. The bandwidth cost at $0.08/GB for cloud egress is $13/day per camera - $1,300/day for 100 cameras. Edge inference processes the video locally and transmits only metadata (detection events, alerts), reducing bandwidth by 99.9%. Similarly, healthcare, financial, and defense applications where data privacy regulations prohibit transmitting raw data to cloud inference endpoints are edge-only deployments by legal requirement.

02

The Jetson Orin Family: AGX, NX, Nano Performance and Pricing

NVIDIA's Jetson Orin family covers the edge AI spectrum from sub-15W embedded modules to 75W full-compute AGX systems. The Jetson AGX Orin 64GB (12-core ARM CPU, 2048 CUDA cores, 64 Tensor Cores, 64GB LPDDR5) delivers 275 TOPS at INT8 precision. At a system cost of approximately $2,000 per unit, it is the most capable edge AI computer available for local inference. The power envelope of 15-60W means it can run continuously in factory environments, vehicles, or field deployments without active cooling in many configurations.

The Jetson Orin NX 16GB module (8-core ARM CPU, 1024 CUDA cores, 32 Tensor Cores, 16GB LPDDR5) delivers 100 TOPS at INT8 within a 10-25W power envelope. The module form factor (69.6 x 45mm) enables integration into embedded systems where space is constrained. The per-unit cost of roughly $700 makes it the most popular choice for mid-range edge AI deployments. It runs YOLOv8s at 200+ FPS, Llama 3.2 3B at 15-20 tokens/second with INT4 quantization, and small vision-language models suitable for industrial inspection.

The Jetson Orin Nano 8GB (6-core ARM CPU, 512 CUDA cores, 16 Tensor Cores, 8GB LPDDR5) delivers 40 TOPS at INT8 within a 7-15W power envelope. At approximately $250 per unit, it addresses the high-volume, cost-sensitive edge AI market. The Nano is capable of running vision models (classification, object detection, segmentation) at real-time frame rates and small LLMs (1B-3B parameters) for simple text classification or instruction following. It cannot run 7B+ models even with aggressive quantization due to memory constraints.

ModelTOPS (INT8)PowerPriceMax LLM Size
AGX Orin 64GB27515-60W~$2,0007-13B (INT4)
Orin NX 16GB10010-25W~$7003-7B (INT4)
Orin Nano 8GB407-15W~$2501-3B (INT4)
03

Optimizing Models for Edge Deployment: Quantization, Pruning, and Distillation

Edge deployment of large language models requires aggressive optimization that is rarely necessary for cloud inference on H100 or B200 GPUs. The Jetson AGX Orin's 64GB LPDDR5 provides roughly 30-35GB of usable memory for model weights after the operating system and runtime reserve. A 7B parameter model at FP16 requires 14GB of memory, leaving margin for KV cache and runtime. But the deeper constraint is compute throughput: the AGX Orin's 275 TOPS at INT8 corresponds to roughly 50-70 tokens/second for a 7B model - adequate for single-stream inference but limiting for high-throughput serving.

INT4 quantization is the enabler for larger models on edge hardware. A 7B model at INT4 occupies 3.5GB of memory, leaving room for larger batch processing and longer context windows. A 13B model at INT4 occupies 6.5GB, fitting comfortably on the AGX Orin 64GB. The accuracy degradation from INT4 quantization varies by model and calibration method. GPTQ and AWQ quantization methods achieve INT4 perplexity within 1-3% of FP16 baseline for most 7B-13B models when using calibration data representative of the deployment domain. The latency improvement from INT4 is roughly 1.5-2x versus INT8 due to reduced memory bandwidth pressure.

Model distillation - training a smaller student model to mimic a larger teacher model - is the highest-effort but highest-return optimization for edge deployment. A distilled 3B model can match the task-specific accuracy of a 7B teacher model for narrow domains (code generation, customer support intent classification, summarization of specific document types) while running at 3-4x the throughput on Jetson hardware. The distillation training run on cloud GPUs (8x H100 at $9.20/hr on ClusterBid) costs $500-2,000 depending on dataset size and training duration. The one-time distillation cost is recovered within weeks through reduced edge hardware requirements or increased throughput per edge device.

04

TensorRT for Jetson: From Model to Optimized Inference Engine

TensorRT is the standard inference optimization toolkit for NVIDIA edge hardware. The optimization pipeline: import the model from PyTorch, ONNX, or TensorFlow, apply graph optimizations (layer fusion, constant folding, operator elimination), quantize to INT8 or FP16 with calibration data, and generate a TensorRT engine optimized for the specific Jetson GPU architecture. The optimization typically improves inference throughput by 2-3x compared to unoptimized PyTorch execution on the same Jetson hardware.

The key TensorRT optimizations that benefit Jetson deployments: kernel auto-tuning selects the best CUDA kernel implementation for each layer given the Jetson GPU's specific SM count and memory configuration. INT8 calibration using a representative dataset reduces the accuracy loss from quantization by identifying the optimal dynamic range for each layer's activations. DLA (Deep Learning Accelerator) offloading moves supported layers to Jetson's dedicated DLA hardware, freeing GPU compute resources for layers that require CUDA. The DLA handles up to 80% of layer types in common vision models, reducing GPU utilization by 40-60%.

TensorRT-LLM extends TensorRT support to large language models on Jetson hardware. It introduces in-flight batching, paged KV cache management, and speculative decoding for edge deployments of LLMs. On Jetson AGX Orin, TensorRT-LLM with INT4 quantization achieves 25-35 tokens/second for Llama 3.2 8B and 40-50 tokens/second for Qwen 2.5 7B. The LLM optimization pipeline requires more calibration data than vision models (typically 1,000-5,000 representative prompts for INT4 calibration) and 30-60 minutes of optimization time on the Jetson device or 5-10 minutes on an H100 GPU (which can generate a Jetson-compatible engine through cross-compilation).

05

Edge Deployment Strategies: On-Device, Hybrid, and Split Inference

On-device inference runs the entire model on the Jetson hardware with no cloud dependency. This is the simplest deployment model and the only option for applications with strict latency, bandwidth, or privacy requirements. The application developer must carefully manage memory: the model weights, runtime engine, application code, and working buffers must fit within the Jetson's available memory (8-64GB depending on the module). On-device inference is suitable for all vision models, small LLMs (up to 3-7B parameters on AGX Orin), and any model that must function without network connectivity.

Hybrid inference runs a small on-device model for initial processing and escalates to cloud GPU inference when the edge model's confidence is low or when a more capable model is required. The pattern: a Jetson-deployed distilled 3B model handles 80-90% of inference requests with adequate accuracy. When the edge model's output confidence falls below 0.9, the request is forwarded to a 70B model running on cloud H100 GPUs. The hybrid approach achieves the user experience quality of cloud inference for 10-20% of requests while running 80-90% at edge latency with zero cloud GPU cost.

Split inference partitions the model between edge and cloud. Typically, the first few layers (feature extraction) run on the Jetson, compressing the raw input into a compact feature vector that is transmitted to the cloud for the remaining model layers. The feature vector is significantly smaller than the raw input (kilobytes versus megabytes for video frames), reducing bandwidth by 100-1,000x versus full cloud inference. Split inference enables cloud-scale models (70B+ parameters) on bandwidth-constrained connections where the full model cannot fit on the edge device but the latency of full cloud inference is unacceptable.

06

Total Cost of Ownership: Edge Jetson vs Cloud H100 for Inference

The TCO comparison between edge and cloud inference depends on deployment scale, inference volume, and latency requirements. A single Jetson AGX Orin at $2,000 hardware cost, running 24/7 at 30W average power ($0.12/kWh), costs approximately $350 per year in hardware amortization and power. If this device handles 10 inference requests per second (an object detection model running at 10 FPS), the annual inference capacity is 315 million inferences at a cost of $1.11 per million - competitive with cloud GPU inference at $0.37 per million tokens (which corresponds to roughly 5-10 million inferences per dollar for LLM calls).

The breakeven point between buying edge devices and renting cloud GPUs varies by workload. For a deployment of 1,000 edge devices running continuously, the hardware investment is $250,000-2,000,000 depending on Jetson module selection. The equivalent cloud GPU capacity (1,000 concurrent inference streams, assume 100 H100 GPUs at $1.15/hr) costs $100,000/month in GPU time and $12,000/month in data transfer for streaming data. The edge deployment breaks even in 2-20 months depending on device tier. After breakeven, the edge deployment operates at roughly 20-30% of the ongoing cloud GPU cost.

The hidden cost of edge AI is maintenance and updates. Cloud GPU inference deployments can update models centrally with a single deployment operation. Edge deployments require over-the-air (OTA) update infrastructure, model validation across diverse hardware configurations, and field support for device failures. The maintenance overhead adds roughly 15-25% to the edge TCO. Teams planning edge AI deployments should budget 0.5-1 FTE per 500-1,000 edge devices for ongoing maintenance, reducing the cost advantage of edge versus cloud but still leaving a net savings for large-scale deployments.

07

Edge AI Roadmap: Jetson Thor and Beyond

NVIDIA's Jetson roadmap includes the Thor platform, expected to deliver approximately 2,000 TOPS at INT8 - a 7x improvement over AGX Orin. Thor will support models up to 70B parameters at INT4 quantization on a single edge module, fundamentally changing the edge AI landscape. The power envelope is expected to be 75-150W, requiring active cooling but still dramatically lower than the 700W of a single H100 GPU. Thor will enable edge deployment of models that currently require multi-GPU cloud inference, opening edge AI applications for large language models in robotics, autonomous systems, and field intelligence.

The convergence of edge and cloud in 2026-2028 is not about one replacing the other but about seamless hybrid architectures where models can run on any compute tier depending on availability and requirements. Edge-first deployment with cloud backup, where inference starts on the local Jetson and seamlessly transitions to cloud GPU when additional compute is needed, is the dominant emerging pattern. This architecture requires runtime systems that can split, migrate, and rejoin inference state between edge and cloud without user-perceptible latency gaps.

For teams evaluating edge AI investments in mid-2026: invest in model optimization and quantization tooling today rather than waiting for more powerful edge hardware. Models that are optimized for Jetson Orin will benefit proportionally from Thor's higher throughput. Develop your deployment infrastructure (OTA updates, device monitoring, analytics) on the current Jetson generation so that the operational foundation is ready when Thor hardware arrives. The software investment in edge AI deployment tools does not become obsolete with new hardware - the runtime systems, model optimization pipelines, and OTA infrastructure transfer directly to newer Jetson generations.

Filed under
Edge AIJetson OrinEmbedded GPUNVIDIA JetsonInference at EdgeTensorRTEdge Deployment