All essays
TechnicalDEEP DIVEFEB 2026

AI Model Serving for Mobile and Edge: On-Device Inference, Model Compression, and GPU Tradeoffs

Infrastructure guide for edge and mobile AI deployment: model compression (quantization, pruning, distillation), on-device GPU and NPU inference, hybrid edge-cloud architectures, and GPU tradeoff analysis for latency-constrained AI workloads.

01

THE EDGE AI COMPUTE SPECTRUM

Edge AI spans a 1,000x compute range: microcontrollers with 256 KB SRAM, mobile NPUs with 10-30 TOPS INT8 (7B models in 4-bit), and edge server GPUs with 50-200 TOPS (Jetson Orin, L4). The decision between edge and cloud is driven by latency (sub-50ms needs edge), data privacy, connectivity reliability, and TCO.

A mobile app running a 7B LLM on-device costs $0.0001-0.0005 per query (amortized) versus $0.001-0.005 for cloud API, making on-device 3-10x cheaper at volume. However, the device hardware cost ($20-50 for edge NPU) and deployment complexity shift the balance toward cloud for low-volume scenarios.

02

MODEL COMPRESSION: QUANTIZATION, PRUNING, AND DISTILLATION

Weight-only INT4 quantization on Llama-3.1-8B reduces memory from 16 GB to 2 GB with 1.5-3 percent MMLU drop. Activation quantization requires SmoothQuant or QuIP for handling outliers, enabling INT8 activation with <1 percent quality loss.

Structured 2:4 sparsity plus INT8 quantization achieves 4-6x compression with 2-4 percent degradation, enabling 7B models on mobile GPUs. Knowledge distillation (Phi-3-mini, TinyLlama) produces purpose-built edge models 40-60 percent smaller with minimal quality loss.

For production edge AI: distilled model + INT4 quantization achieves 5-8x total compression versus the original teacher.

MethodSize ReduxQualityHardwareSpeedupComplexity
INT8 W8A82x<1% MMLUH100/L4/Orin1.8-2.2xLow
INT4 W4A163.5-4x1-3%H100/ANE/L42.5-3.5xMedium
2:4 Sparsity2x<2%Ampere+ GPUs1.5-2xMedium
Distillation2-3x0.5-3%Train GPU2-3xHigh
Combined6-10x3-5%Edge NPU/GPU4-8xVery High
03

MOBILE GPU VERSUS NPU: HARDWARE ACCELERATION TRADEOFFS

Mobile NPUs (Apple Neural Engine, Qualcomm Hexagon) offer 2-3x better power efficiency than GPUs for attention-based models but support fewer operations. Apple's Apple Intelligence uses hybrid deployment: smaller models run entirely on ANE, larger models use GPU fallback for unsupported operations.

Hybrid execution achieves 80-120 tok/s on Apple A17 Pro for a 3B INT4 model versus 40-60 tok/s on GPU-only. The edge server tier (Jetson Orin NX 16GB, 100 TOPS INT8) runs full 7B INT4 models at 30-50 tok/s on 15W TDP, replacing $10,000-50,000 in cloud GPU inference per year.

HardwareTOPS INT8Power7B Speed3B SpeedCostUse Case
A17 Pro ANE352-3WN/A60-90 tok/sIn phoneMobile chat
A17 Pro GPU4 FP165-10W5-15 tok/s40-60 tok/sIn phoneFallback
Snap 8 Gen3 NPU452-4WN/A50-80 tok/sIn phoneAndroid AI
Orin NX 16GB10010-25W30-50 tok/s80-120 tok/s$600-800Edge server
Orin AGX 64GB27515-60W50-80 tok/s120-180 tok/s$2,000Medical
L4 (edge cloud)242 FP870-150W120-180 tok/s280-350 tok/s$3,000Retail AI
04

HYBRID EDGE-CLOUD ARCHITECTURE FOR AI INFERENCE

The practical pattern: device runs a small model for common cases, falls back to cloud GPU for complex queries. A lightweight router (MobileBERT, TinyBERT) predicts query difficulty. For mobile LLM assistants, 60-80 percent of queries are on-device (3B INT4), 20-40 percent route to cloud (70B). This reduces cloud GPU costs by 60-80 percent.

The cloud gateway must handle smaller request sizes, higher latency tolerance, and device-specific authentication. The cloud GPU pool for edge fallback uses batch sizes of 1-4 because fallback requests arrive asynchronously. During network outages, the on-device model serves all queries with 10-15 percent accuracy degradation.

05

DEPLOYMENT FRAMEWORKS AND TOOLCHAINS

Apple's CoreML + ANE deployment requires Xcode model conversion (`coremltools`), operator fallback registration, and on-device profiling. Operator coverage is 70-80 percent for modern architectures. Qualcomm's AI Hub provides QNN SDK with ~90 percent LLM operator coverage and HTP backend for INT4 via AWQ calibration.

For cross-platform deployment, Google's MediaPipe and NVIDIA TensorRT are primary options. TensorRT provides maximum performance on NVIDIA edge hardware with explicit operator fusion and kernel auto-tuning, but requires per-device compilation (30-60 minutes per model variant), precluding runtime model adaptation.

06

B200 AND THE EDGE-CLOUD CONTINUUM

B200 lowers cloud inference costs by 3-5x versus H100. At $0.50-0.80/GPU-hour reserved, cloud cost for a 70B model is $0.001-0.002 per query. For 10M queries/month, this is $10,000-20,000 (B200) versus $30,000-60,000 (H100). This makes cloud-only architectures more viable for mid-scale deployments.

However, network round-trip (20-50 ms wired, 100-500 ms cellular) versus zero network delay for on-device means the user-perceptible difference persists. The long-term architecture is three-tier: on-device NPU for real-time (30 ms), edge server GPU for medium (50-100 ms), and B200 cloud for complex queries (100-200 ms).

Filed under
Edge AI InferenceMobile Model DeploymentModel Compression GPUOn-Device MLNPU vs GPUHybrid Edge Cloud AIQuantized AI Models