THE EDGE AI COMPUTE SPECTRUM
Edge AI spans a 1,000x compute range: microcontrollers with 256 KB SRAM, mobile NPUs with 10-30 TOPS INT8 (7B models in 4-bit), and edge server GPUs with 50-200 TOPS (Jetson Orin, L4). The decision between edge and cloud is driven by latency (sub-50ms needs edge), data privacy, connectivity reliability, and TCO.
A mobile app running a 7B LLM on-device costs $0.0001-0.0005 per query (amortized) versus $0.001-0.005 for cloud API, making on-device 3-10x cheaper at volume. However, the device hardware cost ($20-50 for edge NPU) and deployment complexity shift the balance toward cloud for low-volume scenarios.
MODEL COMPRESSION: QUANTIZATION, PRUNING, AND DISTILLATION
Weight-only INT4 quantization on Llama-3.1-8B reduces memory from 16 GB to 2 GB with 1.5-3 percent MMLU drop. Activation quantization requires SmoothQuant or QuIP for handling outliers, enabling INT8 activation with <1 percent quality loss.
Structured 2:4 sparsity plus INT8 quantization achieves 4-6x compression with 2-4 percent degradation, enabling 7B models on mobile GPUs. Knowledge distillation (Phi-3-mini, TinyLlama) produces purpose-built edge models 40-60 percent smaller with minimal quality loss.
For production edge AI: distilled model + INT4 quantization achieves 5-8x total compression versus the original teacher.
| Method | Size Redux | Quality | Hardware | Speedup | Complexity |
|---|---|---|---|---|---|
| INT8 W8A8 | 2x | <1% MMLU | H100/L4/Orin | 1.8-2.2x | Low |
| INT4 W4A16 | 3.5-4x | 1-3% | H100/ANE/L4 | 2.5-3.5x | Medium |
| 2:4 Sparsity | 2x | <2% | Ampere+ GPUs | 1.5-2x | Medium |
| Distillation | 2-3x | 0.5-3% | Train GPU | 2-3x | High |
| Combined | 6-10x | 3-5% | Edge NPU/GPU | 4-8x | Very High |
MOBILE GPU VERSUS NPU: HARDWARE ACCELERATION TRADEOFFS
Mobile NPUs (Apple Neural Engine, Qualcomm Hexagon) offer 2-3x better power efficiency than GPUs for attention-based models but support fewer operations. Apple's Apple Intelligence uses hybrid deployment: smaller models run entirely on ANE, larger models use GPU fallback for unsupported operations.
Hybrid execution achieves 80-120 tok/s on Apple A17 Pro for a 3B INT4 model versus 40-60 tok/s on GPU-only. The edge server tier (Jetson Orin NX 16GB, 100 TOPS INT8) runs full 7B INT4 models at 30-50 tok/s on 15W TDP, replacing $10,000-50,000 in cloud GPU inference per year.
| Hardware | TOPS INT8 | Power | 7B Speed | 3B Speed | Cost | Use Case |
|---|---|---|---|---|---|---|
| A17 Pro ANE | 35 | 2-3W | N/A | 60-90 tok/s | In phone | Mobile chat |
| A17 Pro GPU | 4 FP16 | 5-10W | 5-15 tok/s | 40-60 tok/s | In phone | Fallback |
| Snap 8 Gen3 NPU | 45 | 2-4W | N/A | 50-80 tok/s | In phone | Android AI |
| Orin NX 16GB | 100 | 10-25W | 30-50 tok/s | 80-120 tok/s | $600-800 | Edge server |
| Orin AGX 64GB | 275 | 15-60W | 50-80 tok/s | 120-180 tok/s | $2,000 | Medical |
| L4 (edge cloud) | 242 FP8 | 70-150W | 120-180 tok/s | 280-350 tok/s | $3,000 | Retail AI |
HYBRID EDGE-CLOUD ARCHITECTURE FOR AI INFERENCE
The practical pattern: device runs a small model for common cases, falls back to cloud GPU for complex queries. A lightweight router (MobileBERT, TinyBERT) predicts query difficulty. For mobile LLM assistants, 60-80 percent of queries are on-device (3B INT4), 20-40 percent route to cloud (70B). This reduces cloud GPU costs by 60-80 percent.
The cloud gateway must handle smaller request sizes, higher latency tolerance, and device-specific authentication. The cloud GPU pool for edge fallback uses batch sizes of 1-4 because fallback requests arrive asynchronously. During network outages, the on-device model serves all queries with 10-15 percent accuracy degradation.
DEPLOYMENT FRAMEWORKS AND TOOLCHAINS
Apple's CoreML + ANE deployment requires Xcode model conversion (`coremltools`), operator fallback registration, and on-device profiling. Operator coverage is 70-80 percent for modern architectures. Qualcomm's AI Hub provides QNN SDK with ~90 percent LLM operator coverage and HTP backend for INT4 via AWQ calibration.
For cross-platform deployment, Google's MediaPipe and NVIDIA TensorRT are primary options. TensorRT provides maximum performance on NVIDIA edge hardware with explicit operator fusion and kernel auto-tuning, but requires per-device compilation (30-60 minutes per model variant), precluding runtime model adaptation.
B200 AND THE EDGE-CLOUD CONTINUUM
B200 lowers cloud inference costs by 3-5x versus H100. At $0.50-0.80/GPU-hour reserved, cloud cost for a 70B model is $0.001-0.002 per query. For 10M queries/month, this is $10,000-20,000 (B200) versus $30,000-60,000 (H100). This makes cloud-only architectures more viable for mid-scale deployments.
However, network round-trip (20-50 ms wired, 100-500 ms cellular) versus zero network delay for on-device means the user-perceptible difference persists. The long-term architecture is three-tier: on-device NPU for real-time (30 ms), edge server GPU for medium (50-100 ms), and B200 cloud for complex queries (100-200 ms).
