GPU METRIC LOGGING DEPTH AND FIDELITY
Experiment tracking tools differ substantially in the depth of GPU metrics they capture natively. Weights & Biases (W&B) offers the most comprehensive GPU telemetry through its `wandb.log` integration with PyTorch: automatically logging GPU utilization, memory allocated, temperature, power draw, and PCIe/NVLink throughput by default when `wandb.init(project="train", config=config)` is called. W&B's NVIDIA DCGM integration pushes per-GPU metrics every 2 seconds during training, including memory bandwidth utilization and SM occupancy. The metrics are visualized in W&B's workspace as time-series panels with per-step and per-epoch aggregation. W&B also surfaces GPU-specific alerts: if GPU utilization drops below 30% for 5 consecutive steps, an automatic alert triggers in Slack or email.
MLflow's GPU logging requires manual instrumentation via `mlflow.log_metric()` or the new `mlflow.autolog()` API. The autolog integration captures basic GPU metrics: `gpu_utilization`, `gpu_memory_used`, and `gpu_temp` but at 30-second polling intervals versus W&B's 2-second. MLflow does not capture NVLink throughput, SM occupancy, or memory bandwidth utilization. Comet sits between them: `comet_ml.Experiment()` with `auto_metric_logging=True` captures GPU metrics at 10-second intervals and exposes GPU topology information (which GPU index, PCIe bus ID, GPU model name). Comet's notebook cell-level GPU tracking is unique: it correlates GPU utilization spikes with specific Jupyter notebook cells, useful for debugging data loading bottlenecks during training development.
| GPU Metric Depth | W&B | MLflow | Comet |
|---|---|---|---|
| GPU Utilization (%) | 2-sec polling | 30-sec polling | 10-sec polling |
| GPU Memory (allocated/free) | Per-step auto | 30-sec auto | 10-sec auto |
| GPU Temperature / Power | Via DCGM | Manual only | 10-sec auto |
| NVLink / PCIe Throughput | Via DCGM | Not captured | Not captured |
| SM Occupancy | Via DCGM | Not captured | Not captured |
| GPU Topology / Bus ID | Auto-detected | Not captured | Auto-detected |
| Training Step Alignment | Yes (step sync) | Partial | Partial |
ARTIFACT STORAGE: CHECKPOINTS, MODELS, AND COST
GPU training runs produce large artifacts: model checkpoints (10-140 GB for 7B-70B parameters), optimizer states (2x weight size), and training logs. The storage cost and bandwidth for these artifacts vary significantly across platforms. W&B Artifacts stores all files in W&B's cloud storage with a free tier of 100 GB and paid tiers at $0.023/GB/month for storage and $0.12/GB for egress. For a 70B model checkpoint at 140 GB, each saved checkpoint costs $3.22/month in storage and $16.80 per egress download. With checkpoints every 500 steps on a 10,000-step training run, the storage cost is $64.40 per month for 20 snapshots, plus $336 for egress on each evaluation download. Teams training multiple models per week can easily exceed $1,000/month in W&B artifact costs.
MLflow's self-hosted artifact store eliminates per-GB costs but incurs infrastructure expense. An MLflow Tracking Server with an S3-compatible artifact store (MinIO or AWS S3) costs $50-150/month in infrastructure for a medium team (5-10 active projects). The artifact storage is billed at S3 rates ($0.023/GB/month standard), identical to W&B's raw storage, but egress is free within the same cloud region. MLflow supports artifact-version deduplication: if two experiments use the same base model file, MLflow stores one copy and creates a symlink, reducing storage by up to 90% for LoRA training workflows where base model checkpoints are identical across runs. Comet's artifact storage charges $0.025/GB/month with 500 GB free for paid plans at $299/month per user. The key cost tradeoff: W&B is more expensive but zero-ops, MLflow is cheaper at scale but requires infrastructure maintenance.
| Cost Dimension | W&B (Team Plan) | MLflow (Self-Hosted) | Comet (Team Plan) |
|---|---|---|---|
| Platform Base Cost | $99/user/month | $50-150/mo infrastructure | $299/user/month |
| Artifact Storage | $0.023/GB/month | $0.023/GB/month (S3) | 500 GB free, then $0.025/GB |
| Artifact Egress | $0.12/GB | Free (same region) | $0.10/GB |
| 70B Checkpoint Storage (20 runs) | $64.40/mo | $64.40/mo (S3) | Included in 500 GB free |
| Multi-Run Dedup | Manual | Automatic (S3 dedup) | Manual |
| Max Artifact Size | 50 GB per file | No limit (S3) | 10 GB per file |
| Egress from Cluster | Internet (slow/expensive) | Same-region (fast/free) | Internet (slow/expensive) |
TRACKING SERVER ARCHITECTURE AND LATENCY
The tracking server's proximity to GPU compute significantly impacts training workflow reliability. W&B's managed cloud service requires each training GPU to send metrics and artifacts over the internet. For a cluster running inside a VPC, each `wandb.log()` call traverses NAT gateway to the public internet, adding 5-50ms latency per call. With 100 log calls per training step and 100 steps per minute, this adds 500-5,000ms per minute of training wall-clock overhead, equivalent to 1-8% of total training time. For 8x H100 clusters at $20/hr total GPU cost, W&B cloud logging adds $0.20-1.60/hr in hidden GPU overhead. W&B offers a private SaaS deployment within AWS/Azure VPCs at 2-3x the base subscription price.
MLflow's self-hosted architecture avoids internet latency entirely. The Tracking Server runs alongside the GPU cluster, either on a dedicated CPU node or as a sidecar container. With MLflow's `default-artifact-root=s3://mlflow-bucket` pointing to an S3 bucket in the same region, artifact uploads complete at 5-10 Gbps versus 50-200 Mbps to W&B's cloud. The latency per `log_metric()` call is 0.5-2ms versus 5-50ms for W&B's cloud API. For a 10,000-step training run with 50 metric logs per step, MLflow self-hosted saves approximately 4.2 hours of wall-clock time versus W&B cloud (50ms vs 1ms per log, 500,000 log calls = 25,000 seconds = 6.9 hours cloud vs 0.5 hours local). Comet's hybrid architecture stores metrics locally and syncs asynchronously to their cloud, offering a middle ground: 2-5ms metric latency with eventual sync to cloud for team collaboration.
DISTRIBUTED HYPERPARAMETER SWEEPS ON GPU CLUSTERS
W&B Sweeps is the most mature distributed hyperparameter optimization system. It launches sweep agents as Kubernetes pods or Slurm jobs, each claiming a set of hyperparameters from the W&B cloud scheduler. For a sweep with 256 trials of Llama 8B LoRA on an 8x H100 cluster, W&B schedules 8 concurrent agents (one per GPU) and manages result aggregation, early stopping, and Bayesian optimization. The sweep database stores all trial configurations, metrics, and artifacts with a searchable interface for comparing runs across sweeps. W&B's early stopping using the `hyperband` scheduler terminates underperforming trials after 30% of training, cutting total GPU consumption by an average of 40-55% for medium-sized sweeps.
MLflow's native hyperparameter tuning is less developed than W&B's but can be composed with Optuna or Ray Tune for distributed sweeps. The integration uses MLflow's Tracking API from each trial: `with mlflow.start_run(nested=True):` for parent-child run relationships. Ray Tune's `Tuner.fit()` can log to MLflow via `Callback(MlflowCallback())`, enabling distributed hyperparameter search across GPU clusters with MLflow as the centralized tracking backend. The advantage of the MLflow+Ray approach is cost: the entire sweep infrastructure runs within the cluster without per-trial cloud API costs. For sweeps with 1,000+ trials, the W&B sweep cost (included in the per-user license) versus self-hosted MLflow infrastructure cost ($50-150/month) makes MLflow substantially cheaper for high-volume hyperparameter search.
PRODUCTION DEPLOYMENT PATTERNS AND CROSS-PLATFORM SYNC
A common production pattern uses both MLflow and W&B: MLflow as the self-hosted artifact and model registry for cost control, and W&B for experiment visualization and team collaboration. The MLflow tracking URI is set as an environment variable `MLFLOW_TRACKING_URI=http://mlflow-server:5000` on all cluster nodes. Training scripts log metrics to both systems: `wandb.log(metrics)` for real-time visualization, and `mlflow.log_metrics(metrics)` for permanent artifact storage. A background sync script periodically exports W&B run data to MLflow for archival. This hybrid approach adds 15% engineering overhead but reduces cloud costs by 60-80% for artifact-heavy workflows.
Comet's differentiator is its notebook integration for exploratory GPU development. When data scientists run training experiments in JupyterLab on GPU instances, Comet automatically logs GPU utilization and cell execution time without any import statements. For teams that split work between notebooks (prototyping) and production scripts (training), Comet provides a unified dashboard that spans both environments. The tradeoff is Comet's higher per-user pricing ($299/user/month) and smaller ecosystem of integrations compared to W&B and MLflow. For GPU infrastructure teams on ClusterBid, the recommendation is MLflow self-hosted for production training pipelines (lower cost, no egress), W&B for research and development teams that need rich visualization and sweeps, and Comet for organizations heavily invested in notebook-based GPU workflows.