The Multi-Agent Paradigm Shift
Multi-agent AI systems represent a fundamentally different compute profile from single-model inference or training. An agent orchestrator dispatches tasks to specialized sub-agents, each running an LLM inference call, with intermediate results flowing between agents through structured message passing. The total compute consumed is not a single model forward pass but a tree of dependent inference calls, where each level of agent decomposition multiplies the total token generation requirement.
A typical production agent system in 2026 uses a router agent (70B-200B parameters) that delegates to 3 to 10 specialized sub-agents (8B-70B parameters each). Each sub-agent might run 2 to 5 reasoning loops before responding. This means a single user request can trigger 10 to 50 discrete LLM inference calls, each generating 500 to 2000 output tokens. The GPU cluster must handle this burst pattern while maintaining per-request latency under 5 seconds to preserve user experience.
Compute Profile of Agent Systems
Agent workloads are prefill-heavy but in a burst pattern. The router agent's prefill processes the full user context, then sub-agent prefills process the router's output plus any tool call results. Because each sub-agent operates on a different context, KV cache reuse across inference calls is limited. This means the cluster must provision for approximately 3 to 5x the prefill throughput of a comparable single-model serving workload at the same request rate.
Memory pressure is driven by two factors: model weight footprint across multiple specialized models, and the per-request KV cache storage from concurrent agent conversations. A system running 5 different fine-tuned 70B models simultaneously requires approximately 5 x 140GB = 700GB of HBM capacity for weights alone, assuming FP16. Using B200 GPUs with 192GB each, this demands at minimum 4 GPUs before accounting for KV cache overhead.
Sizing Methodology for Agent Clusters
The correct sizing methodology starts from agent topology, not model size. Map the agent tree depth and branching factor, then multiply by average tokens per agent inference to get the total token generation per request. Apply the target requests-per-second (RPS) and latency SLO (p99 under 5 seconds). The GPU compute requirement in TFLOPS is: total tokens per second x average model FLOPs-per-token, divided by the target MFU achievable with continuous batching.
A worked example: a 3-level agent tree with branching factor 3, 1000 output tokens per inference, 50 RPS, and a 70B router model requires approximately 2,500 output tokens per second through the router. Using FP8 on B300 at approximately 400 tokens/second per GPU for 70B inference, the router tier needs at least 7 GPUs. Each of the 9 sub-agent nodes uses a 13B model at approximately 2,000 tokens/second per GPU, requiring roughly 23 additional GPUs total. The minimum cluster is 30 GPUs.
| Component | Model Size | Tokens/Sec Per GPU | GPUs Required | HBM per GPU |
|---|---|---|---|---|
| Router Agent | 70B | 400 (FP8, B300) | 7 | 192 GB |
| Reasoning Sub-Agent | 70B | 400 (FP8, B300) | 9 | 192 GB |
| Tool-Use Sub-Agent | 13B | 2,000 (FP8, B300) | 7 | 80 GB |
| Memory Sub-Agent | 13B | 2,000 (FP8, B300) | 4 | 80 GB |
| Code Sub-Agent | 13B | 2,000 (FP8, B300) | 3 | 80 GB |
| Total | 30 | ~6 TB system |
Inference Caching for Agent Workloads
Agent systems present a unique opportunity for semantic KV cache sharing. When multiple agent turns reference the same tool output or the same retrieved context chunk, prefix caching can reuse KV entries across inference calls. NVIDIA Dynamo introduces distributed KV cache across GPU nodes, which is particularly effective for agent workloads where the router's system prompt and tool descriptions are reused across every request.
With Dynamo-style disaggregated caching, the router agent tier's KV cache hit rate can reach 40 to 60 percent in production agent deployments. This reduces effective GPU demand by reducing the prefill computation per request. For the agent cluster example above, a 50 percent KV cache hit rate on the router tier reduces the total GPU requirement from 30 to approximately 22 GPUs, a 27 percent reduction in cluster cost.
Cluster Topology for Agent Orchestration
Agent clusters benefit from a two-tier topology. The inference serving tier uses H200 or B300 GPUs connected by NVLink within each node, with InfiniBand NDR400 or Spectrum-X Ethernet connecting nodes. The orchestration tier runs the agent framework (LangGraph, CrewAI, AutoGen) on CPU nodes with ample DRAM for conversation state storage and tool execution environments.
The critical topology constraint is latency between the orchestrator and the inference tier. Agent frameworks need sub-10 millisecond response times from the inference API to maintain sub-second agent turn-around. This typically requires deploying the orchestration tier on the same local network as the inference GPUs, ideally within the same rack. Cross-region or even cross-AZ orchestration incurs 30-50ms round-trip penalties that make 5-second end-to-end latency impossible at scale.
Cost Projections at Scale
A 30-GPU B300 cluster for the agent topology described above costs approximately $6,000-$7,500 per day at current spot rates through ClusterBid. At 50 RPS sustained over a 24-hour period, that translates to $5.00-$6.25 per 1,000 agentic requests processed. Compared to single-model inference at approximately $0.50-$1.00 per 1,000 requests, the agent markup factor of 5-12x reflects the multiplicative token generation inherent to multi-agent systems.
The cost trajectory improves with scale. At 200 RPS on the same agent topology, continuous batching efficiency on B300 improves GPU utilization from approximately 45 percent to approximately 75 percent, bringing the per-request cost down to $2.50-$3.50 per 1,000 agentic requests. This is why agent deployments below 10 RPS are hard to justify economically today: the overhead of maintaining the full agent topology is not absorbed by low request volumes.
Our Recommendation
Start with a single-agent architecture before scaling to multi-agent. A single 70B agent handling router and tool-call responsibilities with a monolithic system prompt can serve 80 percent of use cases at one-third the GPU cost. Add specialized sub-agents only when the router agent shows measurable quality degradation on specific task types.
When you do scale to multi-agent, reserve GPU capacity through ClusterBid at the level of the agent tier, not the total cluster. Use longer-term reservations for the router tier GPUs (where model changes are infrequent) and shorter-term or spot capacity for sub-agent tiers that may change as you optimize the agent topology. This tiered reservation strategy typically saves 15 to 25 percent versus buying all capacity at a single commitment level.
