Why Edge Inference Exists as a Separate Category from Cloud
A B200 GPU in a datacenter draws 700W, requires liquid cooling, costs $3.92 per hour, and delivers roughly 4.5 petaFLOPs of FP8 tensor throughput. A Jetson AGX Orin draws 15-50W, costs roughly $2,000 as a module, and delivers 275 TOPS at INT8. The gap in raw throughput is roughly 16x. The gap in power consumption is 14-47x. For a robotics application where the entire payload can only carry a 300 Wh battery, or a surveillance camera that needs to run inference from a PoE budget of 25W, the datacenter GPU is not an option at any price. Edge inference is a separate category because the physical constraints of deployment site dictate the hardware, not the other way around.
The edge AI market in mid-2026 spans at least five distinct power classes: sub-5W for wearables and IoT sensors, 5-25W for cameras and drones, 25-75W for robotics and autonomous vehicles, 75-200W for edge servers and AI boxes, and 200-700W for near-edge or micro-datacenter deployments. Each class has different GPU options, different quantization requirements, and different cost structures. Buying the wrong class means either failing to meet latency requirements or burning power budget that was allocated to sensors or actuators.
NVIDIA dominates the high end of edge with the Jetson line (Orin NX, Orin Nano, AGX Orin, and the upcoming Thor). AMD's XDNA NPU and the Versal AI Edge series compete at the sub-25W range. Intel's Meteor Lake and Lunar Lake NPUs cover the sub-15W PC-class edge inference market. Qualcomm's Snapdragon Ride Flex targets automotive specifically. This guide focuses on the NVIDIA Jetson lineup because it is by far the most common platform in the 25-200W edge AI deployments we see across ClusterBid's buyer base, particularly for robotics, security, and autonomous vehicle inference.
The 2026 Jetson Lineup: Orin Nano, NX, AGX, and Thor
The Jetson Orin family spans four modules at mid-2026. The Orin Nano 8 GB at $249 is the entry point, delivering 40 TOPS at INT8 within a 7-15W power envelope. It runs small vision models like ResNet-50 or YOLOv8n at 60-100 FPS but cannot fit a 7B-parameter LLM or a vision-language model with more than 3B parameters. The Orin NX 16 GB at $499 roughly doubles the TOPS to 70-100 depending on the clock configuration and power mode, with a 10-25W envelope. It is the most popular module for camera-based AI, where 25W is the typical PoE++ budget and the 16 GB of unified memory is just enough for quantized 7B models at 4-bit.
The AGX Orin at $1,999 provides 275 TOPS INT8 in a 15-50W envelope with 64 GB of unified memory. This is the module that runs medium-sized models in real-world robotics deployments. It can run Llama 3.2 8B at 4-bit (AWQ) at roughly 12-18 tokens per second, which is usable for interactive robot control but not for conversational latency. The unified memory architecture, where CPU and GPU share a 64 GB pool, eliminates PCIe transfer costs and keeps effective bandwidth at roughly 200 GB/s between the GPU cores and the memory.
Jetson Thor, shipping in evaluation units as of Q2 2026 with production expected Q4 2026, is the first edge GPU built on the Blackwell architecture. Thor targets 750 TOPS at FP8 (up to 1,500 INT8 TOPS) within a 50-150W envelope. It uses the same GB10 die as the DGX Spark desktop but at lower clocks and with a unified 128 GB LPDDR6 memory pool. First benchmarks from early access partners show Thor running Llama 3.1 8B at FP8 at 45-55 tokens per second, and running Llama 3.2 90B at 4-bit (AWQ + KV8) at 6-8 tokens per second, numbers that would have required a full H100 just two years ago. For a full comparison of Jetson platforms against datacenter GPUs, see our small model inference economics guide.
| Module | TOPS (INT8) | Power (W) | Memory | Price | 7B TPS (AWQ 4-bit) |
|---|---|---|---|---|---|
| Orin Nano 8GB | 40 | 7-15 | 8 GB LPDDR5 | $249 | Not viable |
| Orin NX 16GB | 100 | 10-25 | 16 GB LPDDR5 | $499 | 3-5 TPS |
| AGX Orin 64GB | 275 | 15-50 | 64 GB LPDDR5 | $1,999 | 12-18 TPS |
| Jetson Thor | 1,500 | 50-150 | 128 GB LPDDR6 | $3,999 | 45-55 TPS |
Quantization Is the Enabler: Models That Fit vs Models That Do Not
The defining constraint of edge deployment is unified memory capacity. An AGX Orin has 64 GB total shared between CPU and GPU. A Thor has 128 GB. Both numbers include the operating system, the runtime (JetPack, typically consuming 2-3 GB), the application code, intermediate activation memory, and model weights. The memory available for model weights alone on an AGX Orin after OS and runtime overhead is roughly 50-55 GB. At FP16, a Llama 3.1 70B model requires 140 GB just for weights. It does not fit. The same model at 4-bit AWQ compresses to roughly 38 GB. It fits with headroom for activations.
The practical quantization tiers for edge inference in 2026 are clear. FP8 inference works well on Thor (native support, minimal accuracy loss on MLPerf Edge v5 benchmarks) but consumes roughly 8 GB per 7B parameters. INT4 AWQ compresses to approximately 4 GB per 7B parameters with measured accuracy deltas of 0.5-1.2% on common benchmarks (MMLU, GSM8K, HumanEval). INT4 with FP8 KV cache (the KV8 pattern popularized by Llama.cpp and AWQ in early 2026) provides a good middle ground: weights at 4-bit, attention cache at 8-bit, roughly 4.5 GB per 7B model parameters.
The cost-effectiveness of quantization on edge hardware is extreme because the alternative is not a worse model. It is no model at all. A robot that cannot run an LLM on-board must either stream sensor data to a cloud endpoint (adding 50-200 ms of latency per round-trip and requiring a persistent cellular or WiFi connection) or run a much smaller model with worse reasoning capability. At mid-2026 cellular data costs of roughly $2-5 per GB for machine-to-machine plans, even modest token volumes of 500K per day at an 8B model's 7-bit-per-token output cost add roughly $1.75-3.50 per day in data charges alone, which over a 3-year robot life cycle is $1,900-3,800 just for inference data transport. Quantizing from 16-bit to 4-bit is effectively free by comparison.
| Quantization | Model Size (7B) | AGX Orin Max Params | Thor Max Params | Quality vs FP16 |
|---|---|---|---|---|
| FP16 | 14 GB | ~22B (does not fit 70B) | ~70B | Baseline |
| FP8 | 7 GB | ~45B | ~128B | ~0.1% loss |
| INT4 AWQ | 3.8 GB | ~100B | ~256B | ~0.8% loss |
| INT4 + KV8 cache | 4.5 GB | ~80B | ~200B | ~0.6% loss |
Deployment Domain: Robotics, Autonomous Vehicles, Surveillance, Retail
Robotics is the most demanding edge inference domain because latency requirements are simultaneous and tight. A manipulation robot needs object detection at 30+ FPS (<33 ms per frame), grasp planning at <50 ms, and language understanding for task instructions at <500 ms total round-trip. These workloads typically run on a single AGX Orin or Thor in the robot's compute payload. The power budget is 50-100W from battery. Robot deployment teams we work with at ClusterBid source modules directly from NVIDIA distribution partners and occasionally use our marketplace for bulk JetPack-compatible carrier boards and cooling solutions. Total system cost for robot compute is usually $3,000-8,000 including the module, carrier board, cooling, and enclosure, which represents 5-15% of the total robot BOM.
Autonomous vehicles run at a completely different scale. A Level 4 autonomous truck uses 4-8 AGX Orin or Thor modules (or an alternative like the Qualcomm Snapdragon Ride Flex, or a combination), each dedicated to a sensor modality: camera perception, LiDAR processing, radar, planning, and driver monitoring. Total compute power typically runs 800-2,000 TOPS at a system power budget of 200-600W. The cost per vehicle is $5,000-15,000 for compute, and the inference workload runs 24/7 with safety-critical reliability requirements that preclude any cloud dependency. No AV deployment uses cloud inference for perception. Cloud is used only for map updates and rare edge-case teleoperation.
Surveillance and retail inference workloads have looser latency tolerances (100-500 ms) but tighter power and cost constraints. A typical edge AI camera runs on an Orin NX at 25W PoE++, processing 2-4 video streams simultaneously. The per-unit compute cost must stay under $500-800 for the camera system to be price-competitive with traditional non-AI cameras. Retail use cases (shelf monitoring, checkout-free walkthroughs, footfall analytics) deploy 50-500 cameras per store. At those volumes, the module cost alone at $499 per Orin NX multiplied by 200 cameras is $99,800 per store, before carrier boards, cameras, and installation. The price sensitivity is extreme, which is why surveillance and retail edge deployments tend to use smaller models (YOLOv8, ResNet, lightweight ViTs) rather than LLMs.
Inference Frameworks and Deployment Tooling for Jetson
NVIDIA's JetPack SDK (the BSP for Jetson, currently at JetPack 6.2 in mid-2026) provides cuDNN, TensorRT, and the NVIDIA AI Enterprise runtime for edge. TensorRT integration is the most important factor in achieving published TOPS numbers from Jetson silicon. Running a model through PyTorch directly on Jetson without TensorRT optimization typically achieves 20-40% of the theoretical TOPS. Running the same model through the TensorRT pipeline (FP16 or INT8 calibration, layer fusion, memory planning) consistently hits 70-85% of TOPS on both Orin and Thor in our benchmarking.
The TensorRT workflow for edge adds build-time steps that many ML teams are not used to. Model weights must be pre-converted from native PyTorch or JAX checkpoint format to a Jetson-optimized TensorRT plan file, typically during CI/CD rather than at deployment time. Quantization calibration requires a representative dataset of roughly 500-2000 samples. Each ONNX opset version must be validated against the JetPack TensorRT version. Teams unfamiliar with this workflow frequently lose 2-4 weeks during their first edge deployment just to the tooling learning curve, which is a non-trivial cost when the entire deployment timeline is 12-16 weeks.
There are escape hatches. NVIDIA's TAO Toolkit (Train Adapt Optimize, now at v5.3) automates the retraining-to-TensorRT pipeline for common vision model architectures. For LLM deployment, the LlamaEdge project and llama.cpp with CUDA backend (cuBLAS) now have first-class Jetson support as of Q1 2026, providing a TensorRT-free path that achieves roughly 80-85% of the throughput of a full TensorRT pipeline for standard decoder models. Most edge deployment teams we talk to use TAO or TensorRT for vision models and llama.cpp/cuBLAS for LLMs, accepting the marginal throughput loss in exchange for a significantly simpler deployment pipeline.
Total Cost of Edge Inference: Module + Power + Connectivity
The three-year total cost of an edge inference deployment breaks down into hardware amortization, power, and connectivity. The module cost dominates for Orin-class deployments. A single AGX Orin at $1,999 amortized over 3 years adds $0.076 per hour. Power at 40W average draw and $0.12 per kWh adds $0.0048 per hour. If the deployment uses cloud fallback for any part of the inference pipeline, cellular data at 100K tokens per day (roughly 1-2 MB of API traffic) adds $0.06-0.15 per day or about $22-55 per year. The module cost is roughly 5-10x the power cost per hour, which is the inverse of the datacenter GPU ratio where power often exceeds hardware amortization.
Thor changes this equation. At a projected $3,999 module price and 100W average draw, the hardware cost is $0.152 per hour and the power cost is $0.012 per hour. The hardware/power ratio is closer to 12:1, even more skewed toward upfront capital. This matters for deployment budgeting because edge deployments are usually capitalized (purchased upfront) while datacenter deployments are often operationalized (rented per hour). An Orin NX deployment of 500 units at $499 each requires $249,500 upfront. The same capacity in the cloud at roughly 275 TOPS equivalent (roughly 20-30% of a single A100) would cost approximately $0.60-1.00 per hour per instance, or $2.6-4.4 million over 3 years at 24/7 operation, assuming the cloud alternative works at all given latency constraints.
The ClusterBid marketplace does not currently broker edge modules directly, but we frequently advise buyers on total-cost comparisons between edge and cloud inference at the procurement stage. For deployments with predictable workload, no mobility constraint, and latency tolerance above 200 ms, the cloud option usually wins on total cost. For anything that has to move, has premium-latency requirements, or operates without reliable network connectivity, edge wins on technical necessity and the cost comparison is secondary.
| Device | 3-Year Hardware | 3-Year Power Cost | 3-Year Connectivity | 3-Year Total |
|---|---|---|---|---|
| Orin Nano ($249) | $249 | $27 | $165 (cloud fallback) | $441 |
| Orin NX ($499) | $499 | $53 | $165 | $717 |
| AGX Orin ($1,999) | $1,999 | $126 | $165 | $2,290 |
| Jetson Thor ($3,999) | $3,999 | $315 | $165 | $4,479 |
| Cloud (A100 equivalent) | $0 cap-ex | $0 cap-ex | $0 | $26,000-44,000 |
The Edge AI Outlook: What Thor Means for 2026-2028 Deployments
Jetson Thor changes the edge inference calculus more than any previous Jetson generation because it closes the gap to datacenter-class inference on the workloads that matter most for advanced edge AI: large vision-language models, real-time multimodal reasoning, and on-device LLM agents. Running Llama 3.1 70B at 6-8 TPS on a 150W module means a warehouse robot can hold a full conversational context with natural language understanding, object recognition, and path planning on the same compute module without streaming any sensor data to a cloud endpoint. That capability did not exist on any edge platform before 2026.
The practical deployment timeline for Thor is the key question. As of June 2026, NVIDIA has shipped evaluation units to roughly 200 partners. Production modules are expected in Q4 2026 with carrier board availability in early 2027. Early production allocation is largely reserved by automotive OEMs (Toyota, Mercedes, and a Chinese OEM that NVIDIA has not named publicly). Non-automotive buyers (robotics, industrial automation, medical devices) should not expect meaningful Thor supply before Q2 2027. For a deeper look at allocation dynamics across GPU product lines, see our B200 procurement timeline guide.
For teams planning edge deployments in the second half of 2026, the recommendation is straightforward. If your inference workload fits on AGX Orin with INT4 quantization (roughly 50 GB of total model and activation memory), deploy on Orin now and treat Thor as a drop-in upgrade when it ships. The JetPack 6.x runtime and TensorRT plan files are backward compatible across Orin and Thor for the operations that both architectures support. If your workload requires more than 64 GB of memory or 275 TOPS, you have three options: wait for Thor, split the workload across multiple Orin modules (which adds cost and complexity), or reconsider whether cloud or near-edge (micro-datacenter with L40S or L4 GPUs) might work within your latency and connectivity constraints.
