All essays
TechnicalDEEP DIVEFEB 2026

Retail and E-Commerce AI Infrastructure: Recommendations and Computer Vision

How retailers and e-commerce companies deploy GPU infrastructure for recommendation systems, visual search, inventory management, and computer vision at store level. Cost analysis.

01

RETAIL AI: THE SCALE CHALLENGE OF RECOMMENDATION AND VISION

Retail and e-commerce AI infrastructure must handle extreme throughput variance with tight latency budgets. Amazon's recommendation system processes 200 million active users generating hundreds of requests per user per session. Each request requires real-time inference through a deep learning recommendation model (DLRM) with billions of embeddings spanning millions of products. Training these models requires 128-512 H100 GPUs in distributed configurations, while serving requires 500-2,000 inference GPUs to maintain sub-50ms response times during peak shopping events like Black Friday and Prime Day.

The workload splits into offline training and online inference at roughly a 30:70 compute ratio, inverted from the LLM market where training dominates. Offline training of the core recommendation model happens weekly or daily using 128-256 GPUs over 8-24 hours. Online inference runs continuously on clusters designed for high throughput, low batch-size processing. The inference GPU market for retail is substantially larger than the training GPU market because each user session generates 50-200 separate recommendation calls, each requiring a forward pass.

Retail WorkloadGPU ClassScaleTraining CadenceInference SLAs
Recommendation EngineH200 141GB128-512 train / 500-2K serveDaily<50ms p99
Visual SearchB200 180GB32-128 GPUsWeekly<100ms p99
Inventory ForecastingH100 80GB16-64 GPUsHourlyBatch, <5 min
In-Store Computer VisionL40S (edge)10-50 per storeMonthly<30ms per frame
Personalized PricingMI325X 288GB32-128 GPUsReal-time<200ms per calc
02

RECOMMENDATION SYSTEMS: DLRM AND TRANSFORMER-BASED ARCHITECTURES

Modern recommendation systems at Amazon, Shopify, and Instacart have shifted from traditional collaborative filtering to transformer-based architectures that incorporate user behavior sequences, product attributes, and real-time context. Meta's DLRM and the newer DLRM-RL approach embed categorical features through learned embedding tables that contain tens of billions of parameters. An e-commerce DLRM with 50 billion embeddings requires approximately 800GB of GPU memory across 8 H200 GPUs using model parallelism. The embedding lookup itself is a memory-bandwidth-bound operation that benefits disproportionately from H200's 4.8 TB/s HBM3e bandwidth versus H100's 3.35 TB/s.

Inference serving for recommendations uses a distinct architecture from LLM serving. Instead of KV-cache and autoregressive generation, recommendation inference is a single forward pass through a transformer encoder followed by a dot-product scoring step across 1,000-10,000 candidate products. The compute requirement is modest per request (0.5-2 ms on H100), but the request volume is enormous. A top-10 e-commerce site processes 500,000-2,000,000 recommendation requests per second during peak hours. Serving this volume requires 400-1,200 H200 GPUs configured with NVIDIA Triton Inference Server using dynamic batching and concurrent model execution.

04

INVENTORY FORECASTING AND SUPPLY CHAIN OPTIMIZATION

Inventory forecasting at scale uses time-series transformer models that predict demand for millions of SKU-warehouse combinations across 52-week horizons. Walmart processes forecasts for 500 million SKU-location pairs weekly. Each forecast requires a forward pass through a 200M-parameter time-series model processing the past 104 weeks of sales data, promotional calendars, weather data, and macroeconomic indicators. The full forecasting run requires approximately 2,000 H100 GPU-hours per week at a cost of $6,000-$7,000, producing demand predictions that reduce inventory carrying costs by 10-20 percent.

The training infrastructure for inventory models runs in a continuous loop, incorporating new sales data daily. A 64-GPU H100 cluster can retrain the full inventory model in 3-4 hours using automatic mixed precision and distributed data parallelism. The financial return on this GPU investment is substantial: a 10 percent reduction in inventory carrying costs at a retailer with $10 billion in inventory equates to $150-200 million in annual savings from reduced warehousing, spoilage, and financing costs. This makes retail inventory AI one of the highest-ROI GPU investments outside of direct revenue generation.

05

IN-STORE COMPUTER VISION AND THE EDGE GPU DEPLOYMENT

Brick-and-mortar retailers are the fastest-growing segment of edge GPU deployment. Amazon Go's Just Walk Out technology, now licensed to 100+ retailers, uses ceiling-mounted cameras feeding a 3D computer vision pipeline that tracks every item a customer picks up. Each store runs 30-100 cameras at 1080p 30fps, processed by 4-8 L40S GPUs or NVIDIA Jetson AGX Orin modules at the edge. The total GPU compute per store is $15,000-$30,000 in hardware plus $3,000-$6,000 monthly in cloud connectivity for model updates and exception handling.

Shelf monitoring represents a more scalable edge use case. Walmart deploys shelf-scanning robots with NVIDIA Jetson Orin NX modules running real-time object detection to identify empty shelves, misplaced products, and pricing errors across 4,700 US stores. Each robot processes 30-50 shelf images per minute through a YOLOv8 model quantized to INT8, running at 15W power draw. The total edge GPU footprint for Walmart's shelf monitoring program is approximately 5,000 Jetson Orin modules operating at an annual infrastructure cost of $25-40 million, replacing manual shelf checks that cost $150-200 million annually in labor.

Filed under
Retail AIE-Commerce GPURecommendation SystemsVisual SearchComputer Vision RetailInventory AIPersonalization