All essays
GuideGUIDEFEB 2026

AI Inference at the Edge: GPU Options for On-Device LLM Deployment in 2026

Edge GPU options for LLM deployment in 2026: Jetson, RTX, embedded. Model sizes, quantization levels, latency requirements, and real use cases for on-device inference.

01

What 'Edge' Means for AI in 2026

Edge AI inference means running a model on hardware local to the data source rather than on a centralized GPU cluster accessed over a network. The motivation is latency, privacy, bandwidth cost, and off-grid operation. In 2026, the edge GPU landscape has matured: Jetson modules can run 7B-parameter models at usable speeds, and consumer RTX GPUs handle 30B-70B models with quantization.

The constraints are hard. Power budgets range from 15W (Jetson Orin Nano) to 450W (RTX 5090). VRAM caps at 32 GB on consumer cards and 64 GB on Jetson AGX Orin. Memory bandwidth on edge GPUs is 2-10x lower than datacenter hardware. Every deployment requires explicit trade-offs between model quality, latency, power, and cost.

02

Edge GPU Options

NVIDIA dominates the edge GPU market with the Jetson family (Orin Nano, Orin NX, AGX Orin) and the RTX consumer line. The Jetson Orin Nano delivers 40 TOPS INT8 at 15W for under $500. The Orin NX hits 100 TOPS at 35W. The AGX Orin reaches 275 TOPS at 60W with 64 GB of unified memory. The RTX 5090 delivers 1,800 TOPS FP4 at 450W with 32 GB GDDR7.

AMD's embedded offerings (Ryzen AI 300 series with XDNA2 NPU) and Intel's Meteor Lake NPU provide lower-end options for sub-2B parameter models at sub-15W. Apple's M4 Ultra with 128 GB unified memory is an edge-adjacent option for local inference at desktop power levels, delivering roughly 70 tokens/s on a 70B Q4 model.

DevicePowerVRAM / memoryINT8 TOPSPrice range
Jetson Orin Nano15W8 GB40$499
Jetson Orin NX35W16 GB100$699
AGX Orin60W64 GB275$1,999
RTX 5090450W32 GB GDDR7~900$2,000
Apple M4 Ultra~120W128 GB unified~100$5,999
03

Model Sizes Feasible on Edge

The model size you can run on an edge device depends on quantization and available memory. A 7B model in FP16 requires 14 GB. In INT4, it drops to 3.5 GB. A 70B model in INT4 needs 35 GB, plus KV cache overhead for meaningful context lengths. The practical edge deployment sweet spot in 2026 is 7B-30B parameter models at 4-bit quantization.

The AGX Orin with 64 GB unified memory can run Llama 3.1 70B in 4-bit (35 GB weights + ~8 GB KV cache at 32K context) with room to spare. The RTX 5090's 32 GB cap makes 70B Q4 models tight but feasible with 8K context. For devices under 35W, the ceiling is roughly 7B in 4-bit (3.5 GB) or 3B in FP16 (6 GB).

ModelFP16 sizeINT4 sizeMinimum edge GPU
Llama 3.2 3B6 GB1.5 GBOrin Nano
Llama 3.1 8B16 GB4 GBOrin NX
Llama 3.1 70B140 GB35 GBAGX Orin / RTX 5090
DeepSeek-Coder V3 33B66 GB16.5 GBAGX Orin
04

Quantization for Edge Deployment

Quantization is the primary mechanism for fitting large models into edge VRAM budgets. INT4 quantization (4-bit weights) reduces model size by 4x versus FP16 with a typical perplexity increase of 0.5-1.5 points depending on the model and calibration data. INT8 quantization offers a 2x reduction with negligible quality loss for most models.

The key trade-off is that quantization also affects inference speed. INT4 matrix multiplications are faster on hardware with dedicated INT4 tensor cores (RTX 5090, B200) but can be slower on older hardware that emulates INT4 operations in INT8. Jetson Orin supports INT8 natively but lacks hardware INT4, so INT4 on Orin runs at roughly 60% of INT8 throughput. The choice of quantization level must account for both the memory budget and the throughput characteristics of the target hardware.

05

Latency Requirements

Edge inference latency requirements vary by use case. Real-time voice assistants need under 200 milliseconds end-to-end. Interactive chatbots target under 500 milliseconds. Batch processing (document summarization, content moderation) can tolerate 2-5 seconds. These targets determine the minimum viable GPU for a given model.

On an AGX Orin running Llama 3.1 8B at INT8, text generation achieves roughly 45 tokens per second, or 22 milliseconds per token. A 200-token response takes 4.4 seconds, which is acceptable for chat but too slow for real-time voice. On an RTX 5090 running the same model at FP4, throughput exceeds 200 tokens/s, bringing response latency under 1 second. The GPU tier directly determines which use cases are feasible.

06

Bandwidth Constraints

Memory bandwidth on edge GPUs is the binding constraint for token generation speed. The Jetson AGX Orin has roughly 200 GB/s of memory bandwidth (shared between CPU and GPU). The RTX 5090 offers 1.8 TB/s. A 70B model at INT4 generates tokens at roughly 2-3 tokens/s on AGX Orin versus 60-80 tokens/s on RTX 5090. The bandwidth difference is the proximate cause.

Network bandwidth is the other constraint. Edge devices frequently operate on metered or unreliable connections. Downloading a 35 GB model (70B INT4) over a 100 Mbps connection takes 47 minutes. Over a 10 Mbps cellular link, it takes 8 hours. Deployment strategies must account for the initial model download, and incremental update mechanisms (LoRA adapters, delta quantization) are essential for maintaining edge devices in the field.

07

Real-World Edge AI Use Cases

Three edge AI use cases dominate production deployments in 2026. The first is on-device coding assistants running on developer laptops. An RTX 5090 or M4 Ultra running DeepSeek-Coder V3 33B at INT4 provides completion latency under 100 milliseconds without sending code to any cloud service. The second is industrial vision-language models on Jetson AGX Orin for real-time defect detection with natural language querying, replacing fixed-classification models.

The third is privacy-preserving medical document processing on hospital-premises edge servers. A single RTX 6000 Ada or A6000 running Llama 3.1 70B at INT4 can process 50,000 clinical notes per day without any PHI leaving the hospital network. These deployments are growing roughly 3x year-over-year as quantization tools mature and edge GPU costs decline.

Filed under
Edge AIJetsonRTXInferenceQuantizationOn-device LLM