The Utilization Gap Nobody Talks About
The fundamental economics of async batch inference vs real-time GPU serving come down to one number: utilization. Your real-time LLM serving cluster is probably running at 30-45% GPU utilization on average. The GPU sits idle between requests, waiting for the next user to send a message. You're paying full price for 55-70% of idle silicon.
Batch inference flips this entirely. When you're processing a queue of 50,000 documents overnight, or annotating a training dataset, the GPU runs at 85-95% utilization continuously. No waiting. No idle gaps. The same H100 that generates 900 tokens per second at full batch might only deliver 360 effective tokens per second in a real-time API scenario - because of the gaps between requests, request serialization overhead, and the fundamental problem that users don't arrive in neat, perfectly-batched groups.
This isn't a small difference. At $1.03/hr for spot H100 capacity (current ClusterBid marketplace pricing), the batch workload delivers roughly 2.2 billion tokens per dollar. The real-time workload on the same hardware, running on-demand at $2.50/hr, delivers around 250-300 million tokens per dollar. That's a 7-8x cost gap from utilization and pricing combined. If 30-40% of your LLM spend is on workloads that don't require sub-second latency, that's where your next major cost reduction comes from.
Which Workloads Actually Tolerate Async Processing?
The honest answer is more of your workloads than you think. Document analysis pipelines - legal contract review, financial report extraction, medical record summarization - generate their results on a schedule anyway. Nobody needs a 200-page earnings report summarized in under 500 milliseconds. The user uploads, goes to lunch, and expects results within an hour. You're already building async workflows; you're just running them on expensive real-time infrastructure.
Dataset annotation and synthetic data generation are the clearest wins. If you're using an LLM to label 2 million examples for a downstream fine-tuning run, that job is inherently offline. It runs overnight, or over a weekend. The same applies to embedding generation for vector databases - you don't need low latency to compute embeddings for your document corpus, you need throughput. A single H100 node running vLLM in offline batch mode can generate embeddings for roughly 10-15 million document chunks per hour.
The category that surprises people is evaluation. If you run LLM-as-judge pipelines to evaluate model outputs, grade chatbot responses, or run automated red-teaming, those are overwhelmingly async workloads. Teams that already moved their evals to batch report 60-70% cost reductions. The eval job gets queued, processes overnight, and the team reviews results in the morning. Converting eval infrastructure to batch is usually a one-week engineering task with a 6x return.
Why Spot H100 at 80% Utilization Beats Reserved H200 for Batch
The GPU selection question for batch inference is completely different from real-time serving. For a real-time API, you might care about latency per token, time-to-first-token, and the ability to handle request spikes. For batch, you care about throughput per dollar and fault tolerance. These criteria point to different hardware choices.
H100 SXM5 on the spot market currently trades at $0.80-1.50/hr on ClusterBid depending on availability windows, versus $2.50-3.00/hr for on-demand. That 2-3x pricing difference is sustainable for batch because batch workloads can tolerate spot interruptions - you checkpoint your position in the queue and resume. H200 at $3.07-3.16/hr reserved is phenomenal for real-time serving (the 141GB HBM3e means you can fit larger context windows and more concurrent sequences), but for batch, you're paying a 200% premium for latency characteristics you don't need.
The math gets even more interesting when you compare cost-per-token across configurations. An H100 SXM5 running Llama 3.3 70B in FP8 via vLLM offline batch achieves roughly 2,400-2,800 tokens per second at high batch size. At $1.03/hr spot, that works out to approximately $0.10-0.13 per million tokens. The same model on a reserved H200 for real-time serving at 42% average utilization runs at roughly $0.45-0.55 per million tokens. You're paying 4-5x more per token for the latency SLA. For workloads that don't have a latency SLA, that premium is pure waste.
| Configuration | GPU Util | $/hr | Effective Tokens/sec | $/1M Tokens |
|---|---|---|---|---|
| H100 SXM5 spot - batch mode | 88% | $1.03 | 2,464 | $0.12 |
| H100 SXM5 on-demand - real-time | 38% | $2.50 | 912 | $0.76 |
| H200 SXM5 spot - batch mode | 85% | $2.20 | 4,250 | $0.14 |
| H200 SXM5 on-demand - real-time | 42% | $3.07 | 2,100 | $0.41 |
| H100 PCIe spot - batch mode | 90% | $0.80 | 1,800 | $0.12 |
The Cost-Per-Token Math: 90% vs 40% GPU Utilization
The async LLM inference economics for 2026 are straightforward to calculate but the implications are underappreciated. Start with a simple model: you have a single H100 SXM5. Running Llama 3.3 70B at FP8 precision, peak throughput is approximately 2,800 tokens per second at saturated batch size (batch size 256+). At 90% effective utilization in a batch job, you get 2,520 tokens per second. Over an hour, that's 9.07 billion tokens. At $1.03 spot: $0.11 per million tokens.
Now model the same GPU in real-time serving. You handle 500 concurrent users, but at any given second, most GPUs are between requests. The average utilization across a 24-hour period for a production chatbot API hovers between 30-45%. Call it 38%. Effective throughput: 1,064 tokens per second. On-demand pricing: $2.50/hr. Cost per million tokens: $0.65. That's a 5.9x cost difference for exactly the same model, on exactly the same hardware, doing exactly the same computation - just in a different access pattern.
At production scale this compounds. If your team spends $200,000/month on LLM inference and 35% of that ($70,000) is on workloads that could be async (evals, annotation, document processing, report generation), converting those workloads to batch processing on spot H100s could reduce that $70,000 to roughly $12,000-14,000. That's $56,000-58,000/month saved. Annualized: over $650,000. The engineering effort to build the batch pipeline is typically 2-4 weeks. The ROI math is not subtle.
Orchestration Patterns: Ray Data vs vLLM Offline vs NVIDIA Dynamo Batch
Three main patterns dominate batch LLM orchestration in 2026. The first is vLLM offline batch mode, which is the simplest entry point. You load your model once, pass a list of prompts to `LLM.generate()`, and vLLM handles the batching internally with PagedAttention and continuous batching. For jobs under 500K prompts and teams already using vLLM for serving, this is the path of least resistance. The downside: no built-in checkpointing, no distributed work distribution, and the job dies if the node does.
Ray Data is the right choice when you need distributed processing across multiple nodes and fault tolerance. You express your inference pipeline as a Ray Dataset transformation, distribute shards across a cluster of H100 nodes, and Ray handles rescheduling failed tasks automatically. For workloads over a million documents, or anything running on spot instances where interruptions are expected every 6-8 hours, Ray Data is the standard choice. The overhead compared to single-node vLLM is real - maybe 15-20% lower raw throughput due to serialization and coordination - but the fault tolerance and horizontal scaling are worth it for large jobs.
NVIDIA Dynamo batch mode is the newest option and most compelling for teams already running Dynamo for real-time serving. Dynamo 0.3+ introduced dedicated batch queuing that lets you route async work to underutilized prefill nodes during off-peak hours. If you're running a disaggregated prefill/decode cluster and have spare prefill capacity at night, Dynamo batch mode fills that capacity with offline work - effectively getting batch processing at near-zero marginal cost. Teams running hybrid Dynamo deployments report that 20-30% of their overnight batch work runs on what would otherwise be idle prefill capacity.
Building the Hybrid Architecture: Real-Time Cluster Plus Spot Batch Pool
The production pattern that most teams land on is a two-tier GPU cluster: a reserved or on-demand core cluster for real-time serving SLAs, plus a spot-based batch pool that scales with the queue depth. The core cluster handles interactive traffic with guaranteed availability. The batch pool handles everything else - and the key insight is that these two pools can share the same model weights via network-attached storage, eliminating the cold-start penalty of reloading 140GB of weights when a spot node comes online.
For a team processing 50 million documents per month alongside a real-time API, a typical hybrid architecture looks like: 8x H100 SXM5 on-demand for real-time serving (guaranteed capacity, predictable latency), plus a batch pool of 0-32 H100s on spot that scales up each evening and scales to zero on weekends. The batch pool cost is variable: $0/hr when queue is empty, up to $33/hr when running 32 GPUs at full tilt processing overnight queues. Monthly batch costs with this pattern run $2,000-8,000 depending on volume, versus $20,000-30,000 if the same workload ran on the on-demand cluster.
The operational complexity is real but manageable. You need a job queue (Redis, SQS, or Celery all work), a spot interruption handler that checkpoints the current batch position and drains gracefully, and a node pool autoscaler that watches queue depth and spins up/down spot capacity accordingly. Most teams build this in 3-4 weeks. ClusterBid's spot H100 inventory is particularly well-suited here - the interruptible capacity at $1.03-1.50/hr tolerates checkpoint-based recovery and gives you the pricing necessary to make the economics work.
Checkpoint-Based Recovery: Making Batch Jobs Spot-Tolerant
The piece that makes spot GPUs viable for batch inference is checkpoint design. Unlike training checkpoints, inference checkpoints are simple: you just need to track which input IDs have been processed and write completed outputs to durable storage before the node terminates. Spot GPUs on most providers give you a 30-90 second warning before termination. That's enough time to flush in-flight batches, write a checkpoint file to S3 or GCS, and terminate cleanly.
The checkpoint granularity matters. Checkpointing every 1,000 documents means you lose at most 1,000 documents of work on interruption. At 2,500 tokens/document average and 2,400 tokens/second throughput, that's less than 60 seconds of work. The checkpoint write itself takes under a second for a small metadata file tracking processed IDs. When the replacement spot node comes up, it reads the checkpoint, skips processed IDs, and continues from where the previous node left off. End-to-end interruption overhead is typically 2-3 minutes of lost processing per spot reclaim - negligible for overnight jobs.
One pattern worth adopting early: separate output storage from intermediate state. Write completed inference results directly to your final destination (S3 prefix, database, etc.) as each batch completes, rather than accumulating them in node memory. This way, a spot interruption at any point only loses the current in-flight batch. It also means you can fan out results processing to a separate pipeline while inference is still running - your downstream consumers don't wait for the entire job to complete before seeing results.
When Real-Time Serving Is Still the Right Call
Batch inference isn't the answer to every problem. Real-time serving wins when user experience depends directly on latency - conversational AI, autocomplete, code generation where developers wait for results, or agentic systems where each tool call depends on the previous one. The 300ms vs 30-minute distinction is about whether the human is waiting. If they are, batch is wrong. If they're not, you should seriously evaluate whether you need real-time infrastructure.
There's also a secondary consideration around tail latency for interactive workloads. Even if your average document takes 2 seconds to process, a p99 latency of 4 minutes (because your batch queue backed up) is unacceptable for a user-facing feature. Batch is for true async workflows where users receive a notification when results are ready - not for anything where the UI shows a spinner.
The practical heuristic: if your workflow involves a human waiting for a result, use real-time. If your workflow involves a machine generating results that another machine will consume later, use batch. Most enterprise AI stacks have more of the second category than they realize. Run the math on your own workload breakdown before assuming you need to provision for peak real-time throughput across the board.
