CHARGEBACK VS SHOWBACK: WHICH MODEL FITS YOUR CLUSTER?
GPU cluster operators must choose between chargeback (teams pay actual costs from their budgets) and showback (costs reported but not charged). Chargeback creates financial incentives for efficient GPU usage: right-sizing training jobs, cleaning up idle pods, and using preemptible instances. Showback provides visibility without financial friction. Large AI organizations typically use chargeback for production training and showback for experimentation.
The hybrid model splits by partition: production (chargeback, $5.00/GPU-hour) and research (showback, first-come-first-served). Teams with production SLAs buy reserved capacity at 20% discount. Reserved capacity guarantees GPU availability for committed spending: a team committing $200,000/month reserves 2,000 GPU-hours at $3.33/hour.
Cost allocation precision: Level 1 (flat per-GPU-hour rate), Level 2 (differentiate by GPU SKU, add networking and storage surcharges), Level 3 (multi-dimensional rate card tracking power, interconnect, and I/O per job). At 500+ GPU scale, Level 2 provides best cost-to-complexity ratio.
| Model | Financial Impact | Behavioral Incentive | Implementation Cost | Best For |
|---|---|---|---|---|
| Pure Showback | None (informational) | Weak: no budget impact | Low ($5-10K setup) | Small teams, R&D |
| Pure Chargeback | Teams pay full cost | Strong: drives efficiency | High ($50-100K setup) | Production, 500+ GPUs |
| Hybrid (partitioned) | Production pays, research sees | Moderate-strong | Medium ($20-40K setup) | Mature orgs mixed workloads |
| Reserved + On-Demand | Discounted reserved, full on-demand | Strong: encourages commitment | Medium ($30-60K setup) | Large teams predictable demand |
GPU RATE CARD DESIGN AND COST COMPONENTS
Full cost stack per H100 GPU: GPU depreciation (3-year straight-line $30,000 = $1.14/hour), server hardware ($20,000/3yr = $0.76/hour), InfiniBand networking ($0.20/GPU-hour), storage allocation ($0.05/GB-month allocated), power ($0.09/hour), cooling ($0.036/hour), space ($0.02/hour), operations staff ($0.30/GPU-hour). Total: approximately $2.60-3.20/GPU-hour at full utilization.
Rate card mark-up: internal chargeback at cost-plus-10%, internal rate matching cloud pricing to drive adoption, external-facing up to 30-50% margin. H100: $4.50/GPU-hour on-demand, $3.75 reserved (1-month), $1.50 preemptible. Storage: NVMe at $0.15/GB-month, standard NFS at $0.03/GB-month.
Rate changes require 30-day notice, applied only to new jobs (running jobs keep submission-time rate). Rate card published as JSON at an internal URL consumed by the billing system.
| Cost Component | Annual per H100 Node | Hourly per GPU | Notes |
|---|---|---|---|
| GPU depreciation (3yr) | $10,000 | $1.14 | H100 SXM 80GB at $30K MSRP |
| Server hardware (3yr) | $6,667 | $0.76 | CPU + 2TB RAM + 8 NVMe |
| InfiniBand fabric (5yr) | $1,752 | $0.20 | NDR400 switch port amortized |
| Parallel storage | $876 | $0.10 | 4 TB allocated per GPU |
| Power (700W GPU + 200W sys) | $788 | $0.09 | $0.10/kWh, PUE 1.4 |
| Cooling (40% of IT load) | $315 | $0.036 | PUE 1.4 |
| Cluster ops team | $2,627 | $0.30 | 5 FTE / 2000 GPU cluster |
| Total All-In Cost | $22,525 | $2.57 | At 100% utilization |
| Market Comparison (cloud) | $43,800 | $5.00 | AWS p5.48xlarge on-demand |
UTILIZATION TRACKING AND ACCOUNTING INTEGRATION
Slurm provides GPU accounting via sacct --format=AllocTRES. Kubernetes uses Kubecost tracking GPU resource consumption per namespace. The cost allocation pipeline extracts GPU-seconds from the scheduler daily, multiplies by rate card price, and writes to the billing database.
Billing schema: gpu_usage_fact(job_id, team_id, project_id, gpu_sku, gpu_count, start, end, wall_seconds, gpu_seconds, rate, total_cost). Monthly reconciliation aggregates cost by team, comparing against budget. Teams at 80% receive warnings; those at 100% require manager approval.
Idle GPU detection: GPUs allocated but not utilized (<5% for 15+ minutes) at lower rate. Queries Prometheus per allocated GPU. Publishing idle cost per team reduces idle time by 30-50% within 2-3 months.
COST OPTIMIZATION STRATEGIES FOR MULTI-TEAM CLUSTERS
Four levers: Lever 1 (Utilization Improvement): increase from 40-60% to 70-85% through preemptible scheduling and backfill. Lever 2 (Idle Pod Cleanup): auto-terminate pods with GPU utilization < 1% for 15 minutes.
Lever 3 (Reserved Pricing): 20-30% discount for minimum monthly GPU-hour commitment. Lever 4 (Preemptible Partitions): 50-70% discount for research workloads with automatic 5-minute checkpointing.
Weekly cost report emailed to each team: GPU-hours, cost by project, average utilization, idle cost. Teams below 50% utilization receive optimization consultation. Reduces cluster-wide waste by 25-40% within 6 months.
| Optimization Lever | Impact on Cost | Effort | Team Impact | ROI Timeline |
|---|---|---|---|---|
| Increase utilization (40% to 75%) | 50-60% more compute/$ | Medium | Low | 1-3 months |
| Idle pod auto-termination | 10-15% waste reduction | Low | Low | 1-2 weeks |
| Reserved discount | 20-30% discount | Low | Medium | 1 month |
| Right-sizing | 15-25% GPU-hour reduction | Medium | Low | 2-3 months |
| Preemptible partitions | 50-70% discount | Medium | Medium | 1-3 months |
BILLING RECONCILIATION AND AUDIT TRAIL
Monthly reconciliation compares allocation system output against actual costs. Target variance: +/-5% at cluster level. Negative variance triggers rate card review. Positive variance funds upgrades or buffers against price volatility.
Audit trail: immutable entries in billing_audit table with rate card versioning. Teams access self-service portal showing daily GPU costs by project with drill-down to individual job costs including GPU serial numbers.
Internal SLAs: reconciliation within 48 hours of month-end, disputes resolved within 5 business days. >10% billing corrections logged as incidents. Annual external audit verifies rate card calculations.
