WHY MODEL EVALUATION IS A DISTINCT GPU WORKLOAD
Model evaluation runs thousands to millions of independent inference requests with no production latency SLA. The workload is throughput-maximized using batch sizes of 64-256. HELM runs 42 scenarios and 250K-500K queries per model version. A full HELM run on a 70B model costs $5,000-15,000 in GPU compute.
Dedicated evaluation runners achieve 2-3x higher throughput than production inference engines because they use extreme batch sizes, fused kernels optimized for variable-length sequences, and skip production serving overhead (queuing, rate limiting, auth).
EVALUATION FRAMEWORK INFRASTRUCTURE: LM EVAL HARNESS AND HELM
EleutherAI's LM Evaluation Harness supports 300+ tasks. Each task defines few-shot context and prompt template. Adaptive batching groups prompts by length for efficient processing. On H100 with a 70B model, the harness achieves 80-120 tok/s throughput.
HELM uses a distributed runner on Apache Beam, partitioning tasks across GPU workers. Full HELM on 70B requires 40-80 GPU-hours on H100. With 8-16 parallel GPUs, a full evaluation completes in 3-10 hours.
Cost optimization: for multiple-choice tasks, speculative evaluation computes logprobs only for answer tokens (60-80 percent compute reduction). For generative tasks, early stopping when correct intermediate steps appear reduces generation length by 30-50 percent.
| Suite | Tasks | Queries/Run | GPU-Hours 70B | GPU-Hours 7B | Dominant Compute |
|---|---|---|---|---|---|
| LM Eval (full) | 300+ | 100K-200K | 10-25 | 1-3 | Generation 70% |
| HELM (full) | 42 scenarios | 250K-500K | 40-80 | 4-8 | Generation 65% |
| MMLU (5-shot) | 57 subjects | 14,000 | 2-4 | 0.2-0.5 | Logprobs 90% |
| HumanEval+MBPP | 2 sets | 300-500 | 1-2 | 0.1-0.3 | Generation 100% |
| Safety Bench | 100K+ | 100K-300K | 10-30 | 1-3 | Generation 80% |
| Red-Teaming | 1M+ probes | 1M-5M | 40-150 | 4-15 | Generation 85% |
RED-TEAMING AT SCALE: ADVERSARIAL EVALUATION INFRASTRUCTURE
Red-teaming infrastructure has three stages: prompt generation (attacker model generating adversarial variants), probe execution (running each prompt against the target model), and harm classification (evaluating responses). The attacker model is typically a smaller, uncensored model generating thousands of prompt variants via permutation, jailbreak encoding, and role-playing.
Probe execution (100K-1M prompts) dominates GPU cost. On 8x H100 with a 70B target model and vLLM continuous batching, 1M probes complete in 3-6 hours. The harm classifier (fine-tuned DeBERTa-large, 300M) runs on CPU or L4, processing 10,000-50,000 responses per second.
The output feeds a dashboard tracking attack success rate by category. If a category exceeds 2 standard deviations above baseline, an alert triggers model update review. The GPU infrastructure must support evaluation turnaround under 12 hours to keep pace with model development.
SAFETY EVALUATION PIPELINES: FROM COMMIT TO DASHBOARD
A production safety pipeline runs automatically on every model checkpoint and generates a safety report within 2-8 hours. Stages: checkpoint download (10-30 min for 70B), model loading (5-15 min), suite execution (1-6 hours), metric computation (5-30 min), and report generation. The report includes per-category violation rates, regression comparisons, and a pass/fail verdict on safety gates.
The pipeline is orchestrated by a workflow engine (Airflow, Prefect, or Argo Workflows) that manages GPU allocation, retries on transient failures (GPU OOM is common for long-running evaluations), and tracks lineage between model checkpoints and their evaluation results.
Key infrastructure requirement: the evaluation cluster must be isolated from training and production to avoid interference. Dedicated evaluation GPUs (typically 8-16 H100s) are provisioned per evaluation team, with preemptible fallback to training GPUs during idle periods.
| Pipeline Stage | Duration | GPU Type | Autoscaling | Failure Mode |
|---|---|---|---|---|
| Download weights | 10-30 min | CPU/Network | N/A | Network timeout |
| Model load | 5-15 min | 8 H100 | Static pool | OOM on load |
| Benchmark exec | 1-6 hours | 8 H100 | Fixed | GPU OOM mid-run |
| Metric compute | 5-30 min | CPU | K8s HPA | Out of memory |
| Report gen | 2-5 min | CPU | N/A | DB connection |
CONTINUOUS MONITORING AND PRODUCTION EVALUATION
Production model monitoring evaluates every N-th request (typically 1 in 1000) against a held-out evaluation set. The shadow evaluation runs the same request through a reference model (previous version, smaller model, or known-good baseline) and compares outputs. Metrics tracked: output quality (F1, BLEU, ROUGE for relevant tasks), latency percentiles, and embedding drift.
The shadow evaluation GPU cost is 0.1-0.5 percent of production inference cost (sampling 1 in 1000 requests). For a deployment at $100/hour GPU cost, monitoring adds $0.10-0.50/hour. The evaluation pipeline runs on a separate small GPU pool (2-4 L40S) that processes the sampled requests asynchronously.
Embedding drift detection computes the KL divergence between production embedding distribution and a reference distribution daily. A KL divergence threshold triggers automated retraining or rollback. This requires a daily batch embedding job on a single GPU running for 10-30 minutes, costing approximately $1-3 per day.
B200 AND THE FUTURE OF EVALUATION INFRASTRUCTURE
B200's larger memory enables running full evaluation suites for 200B+ parameter models that currently require complex multi-node parallelism. A full HELM suite on a 200B model currently takes 80-160 GPU-hours on H100 (tensor parallelism across 8 GPUs reduces throughput by 30 percent). On B200 with 2x TP, throughput improves by 50-60 percent and GPU-hour cost drops proportionally.
The B200 also enables real-time evaluation during training. Currently, evaluation runs between training checkpoints, creating a 3-10 hour gap between checkpoint and evaluation results. With B200-based evaluation that can parallelize evaluation across model shards more efficiently, turnaround time can drop to 30-60 minutes, enabling evaluation-driven training where model training stops automatically when evaluation metrics saturate.
