WHY GPU MODELS NEED A REGISTRY
A model version comprises multiple artifacts: weights (5-300 GB), tokenizer, inference engine + CUDA version, quantization config, and pre/post pipelines. Without a registry, a version change in any component can cause undetected output differences.
A major API provider's incident in 2025 was caused by deploying the same weights with a newer TensorRT-LLM that changed KV cache layout, undetected for 8 hours because only weight hash was checked.
| Component | Size (70B) | Change Freq | Registry Must Track | Mismatch Impact |
|---|---|---|---|---|
| Weights BF16 | 140 GB | Weekly | SHA-256, precision | Different outputs |
| Tokenizer | 1-10 MB | Rarely | Vocabulary hash | Broken tokenization |
| Inference engine | 2-5 GB | Monthly | CUDA/cuDNN version | KV cache mismatch |
| Quant config | 10-100 KB | Per version | Scale factors | Accuracy degradation |
REGISTRY INFRASTRUCTURE
Three tiers: PostgreSQL for metadata, S3 for artifacts (versioned buckets), local NVMe for hot cache. MLflow extended with S3-compatible storage and parallel upload (8-16 streams) reduces 140 GB model upload from 12 min to 75 sec.
Cross-region replication: replicating 140 GB us-east-1 to eu-west-1 costs $3.50 and takes 2-4 minutes. Monthly sync cost: $100-250 for frequent deployers.
VERSIONING AND LINEAGE TRACKING
Semantic versioning: MAJOR for architecture changes, MINOR for retraining, PATCH for infrastructure changes. Each version stores training run ID, dataset version, base model, and hardware config.
Lineage enables EU AI Act compliance and cost attribution. One enterprise found a PATCH update increased memory by 15 percent, raising monthly costs $23,000. Registry allowed rapid rollback.
| Version Component | Example | Trigger | GPU Cost Impact |
|---|---|---|---|
| MAJOR bump | v3->v4 | Architecture 8B->70B | +5-10x memory |
| MINOR bump | v3.1->v3.2 | New training run | +-5% throughput |
| PATCH bump | v3.1.2->v3.1.3 | TensorRT-LLM upgrade | +-3% latency |
| Quant change | FP16->INT8 | Optimization | -50% memory -2% accuracy |
DEPLOYMENT PIPELINE INTEGRATION
CI/CD pipeline: upload, evaluate, approve, stage, canary at 5%, full rollout. Registry enforces constraints: prevent deploy if eval fails, block rollback to vulnerable versions, require MAJOR approval.
GPU clusters register model cache status. On deploy, registry coordinates pre-fetch to reduce cold-start from 45 to under 3 seconds.
