Why Standardized Scoring
GPU providers advertise peak TFLOPS, HBM bandwidth, and node count. These specs describe the hardware, not the service. Network congestion from noisy neighbors, degraded NVLink links, thermal throttling, and slow storage backends all reduce effective throughput. A standardized scoring methodology isolates these variables by running the same benchmarks on every provider under identical conditions.
ClusterBid defines five scoring axes: Price (cost per effective FLOP), Compute (training throughput), Network (multi-node NCCL bandwidth), Reliability (uptime and interrupt frequency), and Support (ticket response and RMA timelines). Each axis is normalized to a 0-100 scale and weighted according to the target workload profile.
Price Axis: Cost Per Effective FLOP
Price scoring starts with raw hourly rate but adjusts for real-world utilization. A provider charging $3.00/GPU/hr with 95 percent training utilization is cheaper than one charging $2.50/GPU/hr with 60 percent utilization. Effective cost equals hourly rate divided by the fraction of wall time spent in compute. Measure utilization with `nvidia-smi` GPU utilization averaging at least 30 seconds of NCCL training.
The second price adjustment accounts for hidden costs: inter-node bandwidth overage fees, persistent disk charges for checkpoint storage, and minimum commit durations. A provider advertising $2.80/GPU/hr may require a 30-day minimum that drives effective cost to $3.40/GPU/hr when the workload runs for 7 days.
Compute Axis: Training Throughput
Compute scoring runs a standardized GPT-3 175B proxy training job using Megatron-LM with tensor parallelism degree 8 and pipeline parallelism degree 4. The metric is tokens per second per GPU averaged over 1000 steps after a 200-step warmup. This captures real-world training throughput including communication overhead, not just peak TFLOPS.
The benchmark runs at three scales: single node (8 GPUs), four nodes (32 GPUs), and sixteen nodes (128 GPUs). The throughput degradation curve from single-node to sixteen-node reveals how well the provider's interconnect fabric scales. A provider losing more than 15 percent efficiency per doublings of node count receives a reduced score.
| Scale | NCCL All-Reduce Benchmark | Training Throughput |
|---|---|---|
| 1 node (8 GPUs) | 65 GB/s per GPU | 1850 tokens/s/GPU |
| 4 nodes (32 GPUs) | 52 GB/s per GPU | 1420 tokens/s/GPU |
| 16 nodes (128 GPUs) | 38 GB/s per GPU | 940 tokens/s/GPU |
| 64 nodes (512 GPUs) | 22 GB/s per GPU | 510 tokens/s/GPU |
Network Axis: Multi-Node Bandwidth
Network scoring uses NCCL's `all_reduce_perf` benchmark with message sizes from 128KB to 256MB. The bus bandwidth metric at 256MB message size is the primary score component. Providers with InfiniBand NDR400 achieve 320-380 GB/s per node on 8-GPU configurations. Providers using RoCEv2 without congestion control typically deliver 180-240 GB/s.
The second network score component measures variance. Running `all_reduce_perf` ten times and computing the coefficient of variation in bus bandwidth isolates congestion from shared fabric backends. A CV above 0.10 indicates oversubscribed top-of-rack switches or degraded optical transceivers.
Reliability Axis: Uptime and Preemption
Reliability scoring measures observed uptime over a 14-day continuous benchmark. The metric is the fraction of wall time where the GPU cluster is available for training without preemption or hardware failure. Spot instances that are preempted more than once per 48 hours receive a heavy penalty. On-demand nodes with hardware failures that take over 4 hours to replace are penalized proportionally.
The second reliability factor is NCCL timeout frequency. Training runs with NCCL timeout set to 30 seconds. Any NCCL timeout error that is not user-caused (e.g., no kernel launch failures, no out-of-memory events) counts against the provider. More than three NCCL timeouts per 1000 steps results in a 20-point score reduction.
Support Axis: Response and RMA
Support scoring measures first-response time for critical tickets (P0: cluster down) and standard tickets (P2: degraded performance). The benchmark triggers a P0 ticket by intentionally taking a GPU offline via IPMI and measuring the minutes until the provider detects and acknowledges the failure. Providers with automated monitoring score higher.
RMA timelines are scored by requesting a GPU replacement for a reported failure and measuring the hours until a replacement node is provisioned. Advanced RMA providers that pre-stage replacement nodes score 90-100. Providers requiring 48+ hours for cross-ship of replacement GPUs score below 50.
Workload-Specific Weighting
The five axes are weighted differently per workload profile. Training-heavy workloads weight Compute at 35 percent and Network at 25 percent. Inference workloads weight Compute at 25 percent and Reliability at 35 percent because inference cannot tolerate preemption. Development and fine-tuning weights Price highest at 40 percent.
ClusterBid publishes aggregate scores for every provider on the platform, computed from the same reproducible methodology. Users can override weights to match their specific workload profile and see provider rankings update in real time.
| Axis | Training Weight | Inference Weight | Dev Weight |
|---|---|---|---|
| Price | 15% | 15% | 40% |
| Compute | 35% | 25% | 20% |
| Network | 25% | 10% | 10% |
| Reliability | 15% | 35% | 20% |
| Support | 10% | 15% | 10% |
Running Your Own Score
Run the open-source ClusterBid benchmark suite against any provider. The suite deploys a standard PyTorch container, runs `all_reduce_perf`, executes the GPT-3 proxy training loop, and records all metrics in a JSON report. The entire test takes approximately 4 hours on an 8-GPU node.
Compare results against the provider's advertised specs. A provider advertising 4.8 TB/s HBM bandwidth that delivers only 3.2 TB/s under real workloads should raise flags. Share benchmarks with the provider to get SLA adjustments or pricing concessions.
