All essays
TechnicalDEEP DIVEFEB 2026

Test-Time Compute Scaling: Infrastructure Sizing for Reasoning Models

Infrastructure sizing for reasoning models. KV cache requirements for long reasoning chains. GPU concurrency planning. Cost per reasoning query at mid-2026.

01

The Reasoning Compute Paradox: More Thinking Means More GPUs

Test-time compute scaling - the practice of allocating more inference compute to generate better outputs - has emerged as the defining infrastructure challenge of 2026. Models like DeepSeek-R1, Qwen3-Think, and GPT-5 reasoning variants can spend 10,000-50,000 tokens on internal reasoning chains before producing a final answer. A single reasoning query may consume 50-500x more GPU compute than a standard query of equivalent output length. This is not a bug - it is the mechanism by which these models achieve state-of-the-art results on math, coding, and scientific reasoning benchmarks. But it creates an infrastructure cost profile that traditional LLM serving architectures are not designed to handle.

A standard chat query with 1,024-token prompt and 256-token output costs roughly $0.003 at $1.15/hr H100 on-demand ($9.20/hr for 8x H100 node at 100 concurrent queries). A DeepSeek-R1 reasoning query processing the same prompt with 20,000 reasoning tokens (internal chain-of-thought) and a 256-token final answer costs approximately $0.045 - 15x more GPU time per query. For a production system handling 10,000 queries per day with a mix of 30% reasoning queries, the daily GPU cost jumps from $30 to approximately $165 - a 5.5x increase driven entirely by the reasoning token overhead.

The infrastructure implication: reasoning models fundamentally change the GPU sizing equation for inference serving. A cluster sized for standard chat throughput may need 3-10x more GPU capacity to handle equivalent request rates with reasoning models. The KV cache management strategy that works for short-generation models breaks down when generations are 10,000+ tokens long. And the cost per query math that justified GPU rental decisions for standard inference needs to be completely re-evaluated for reasoning workloads.

02

Reasoning Token Generation: Why It Costs More Per Output Token

The cost asymmetry between prefill and decode for reasoning models is more extreme than for standard models. A DeepSeek-R1-style reasoning chain involves the model generating thousands of intermediate reasoning tokens, each produced one at a time in the decode phase. The decode phase on an H100 SXM5 generates approximately 50-80 tokens/second for a 70B model at FP8 with batch size 1. A 20,000-token reasoning chain takes 250-400 seconds of decode time per query. During this period, the GPU is fully occupied with this single request, unable to serve other queries.

Batching is less effective for reasoning models than for standard chat models. Reasoning chains have variable lengths: some queries generate 500 reasoning tokens, others 50,000. The reasoning tokens are generated in the decode phase, where batch efficiency depends on the batch having similar KV cache lengths. In a batch of 8 reasoning queries with varying reasoning lengths, the shorter queries finish early, and the GPU resources are wasted on padding or early termination overhead. Dynamic batching with continuous beam search for reasoning paths can improve efficiency, but most production deployments see batch utilization of 40-60% for reasoning workloads versus 70-85% for standard chat workloads.

The speculative decoding techniques that accelerate standard inference - using a small draft model to propose tokens that the large model verifies in parallel - are less effective for reasoning chains. Reasoner models produce tokens that are contingent on the preceding reasoning tokens, making them harder to predict with a smaller draft model. DeepSeek-R1 with speculative decoding achieves approximately 1.2-1.4x speedup versus 2.0-3.0x for standard Llama 3.1 70B generation. The advantage of parallel verification is reduced when the generation is less predictable.

03

KV Cache Requirements for Reasoning Chains: Memory at Scale

The KV cache is the binding resource constraint for reasoning model serving. Each reasoning token generates one new KV cache entry per attention layer. For a 72B model with 32 layers, 32 attention heads, and 128 dimensions per head at FP8, each token requires 2 x 32 x 32 x 128 = 262,144 bytes (256KB) of KV cache. A 20,000-token reasoning chain produces approximately 5.1GB of KV cache. A 50,000-token chain produces 12.8GB. For 8 concurrent reasoning queries with average 20,000-token chains, the total KV cache memory is approximately 41GB - more than half the available HBM on an H100 SXM5 (80GB).

Memory management strategies that work for standard inference fail at reasoning scale. PagedAttention (vLLM) manages KV cache fragmentation by allocating memory in fixed-size pages. But reasoning chains generate tokens sequentially, and the KV cache grows monotonically - there is no opportunity for cross-request KV cache reuse because each reasoning chain is unique. The effective memory utilization with PagedAttention for long reasoning chains is approximately 70-80%, compared to 85-95% for standard chat workloads, because the page boundaries introduce internal fragmentation that accumulates over long generations.

KV cache offloading to CPU memory is an emerging mitigation for reasoning workloads. The reasoning chain's KV cache from earlier layers (layers 1-16) can be offloaded to CPU pinned memory while the generation continues using layers 17-32. When the backward pass in the attention mechanism needs the earlier KV cache entries, they are prefetched from CPU memory. This technique, implemented in SGLang's reasoning serving branch, enables serving 4x concurrent reasoning queries on an H100 SXM5 versus 2x without offloading. The tradeoff: offloading adds approximately 2-5ms per token of latency for the prefetch step, increasing total reasoning time by 10-15% for 20,000-token chains.

MetricStandard Chat (256 tok)Reasoning (20K tok)Reasoning (50K tok)
KV Cache per Query~65 MB~5.1 GB~12.8 GB
Decode Time (1x H100)~3-5 seconds~250-400 seconds~625-1,000 seconds
Concurrent Q per H100~8-12~2-3~1
Cost Per Query (H100)~$0.003~$0.08-0.12~$0.20-0.32
Batch Efficiency~70-85%~40-60%~30-45%
Speculative Decode Speedup2.0-3.0x1.2-1.4x1.1-1.2x
04

GPU Concurrency Planning for Reasoning Workloads

Concurrency planning for reasoning models starts with the expected number of concurrent reasoning queries at peak. Each concurrent reasoning query on a 70B model at FP8 occupies one H100 SXM5 GPU for 250-400 seconds during the reasoning phase. If your production system expects 10 concurrent reasoning queries at peak, you need at minimum 10 H100 GPUs dedicated to reasoning decode, with an additional allocation for prefill and non-reasoning queries. This is dramatically more GPU capacity than the 2-4 GPUs that would handle 10 concurrent standard chat queries.

The concurrency dimension that teams miss: reasoning chain length variation. If the average reasoning chain is 20,000 tokens but some queries generate 50,000+ tokens (tail ratio of 2.5x), serving must either reserve GPU capacity for the tail requests or implement a timeout mechanism that gracefully truncates reasoning chains that exceed a configured limit. Most production deployments set a maximum reasoning token limit (typically 32,000-48,000 tokens) and instruct the model to produce a final answer when the limit is reached. This truncation reduces tail latency variance at the cost of potential quality degradation for the hardest reasoning problems.

The scheduling strategy for reasoning GPUs differs from standard inference. Reasoning GPUs should use an exclusive-query scheduling model: each GPU processes exactly one reasoning query at a time during the decode phase, rather than the continuous batching used for short-generation workloads. This eliminates the batching inefficiency problem (mismatched KV cache lengths) at the cost of lower theoretical utilization. The utilization optimization comes from sizing the reasoning GPU pool to match the query arrival rate, not from multi-query batching on individual GPUs. A cluster of 32 H100 GPUs handling reasoning decode can process approximately 8-14 concurrent queries (assuming 3-4 GPUs per query for pipeline parallelism) with near-100% GPU utilization during the reasoning phase.

05

Cost Per Reasoning Query: The Mid-2026 Pricing Model

The cost per reasoning query is dominated by the GPU time for the reasoning chain generation. Using ClusterBid's mid-2026 H100 SXM5 pricing at $1.15/hr on-demand ($0.000319 per second): a 20,000-token reasoning chain taking 300 seconds on a single H100 costs $0.096 per query. A 50,000-token chain at 750 seconds costs $0.239 per query. Adding prefill time (~0.5 seconds for a 2K prompt, negligible) and the final answer generation (~3-5 seconds), the total cost is $0.10-0.25 per reasoning query. Compare this to approximately $0.003 for a standard chat query of comparable prompt length - a 30-80x cost multiplier for reasoning.

The economics of reasoning model deployment depend heavily on the acceptance that query throughput is not the optimization metric. A single H100 GPU serving reasoning queries handles approximately 3-4 queries per hour (at 20,000 reasoning tokens average). At $1.15/hr, the GPU cost per query is $0.29-0.38. If your reasoning query throughput requirement is 100 queries per hour, you need 25-35 H100 GPUs at a cost of $28.75-40.25/hr. This is viable for B2B applications (legal document analysis, scientific research, complex code generation) where per-query revenue exceeds $1-5, but it is prohibitive for consumer-scale deployment at current pricing.

The path to lower cost per reasoning query in 2026 is threefold. First, model-level optimizations: DeepSeek-R1's Mixture-of-Experts architecture activates only 37B of 671B total parameters per token, reducing compute per reasoning token by approximately 5x versus a dense 70B model of equivalent reasoning capability. Second, hardware-level optimization: B200's FP4 support and 8 TB/s HBM3e cut reasoning token generation time by approximately 2.5-3x versus H100. Third, algorithm-level optimization: early exiting on confidence for reasoning chains, where the model stops reasoning when it reaches a confidence threshold, reduces average reasoning tokens by 30-50% with minimal quality degradation on benchmarks like MATH-500 and GPQA.

06

Infrastructure Patterns for Production Reasoning Deployments

The recommended architecture for production reasoning serving in mid-2026 uses a three-tier GPU pool: a small prefill pool (B200 or H200 GPUs optimized for compute-intensive prompt processing), a large reasoning decode pool (H100 SXM5 GPUs dedicated to the reasoning chain decode phase), and a final-answer pool (smaller H100 PCIe or even L40S GPUs for the final output generation). The prefill GPU processes the prompt and computes the initial KV cache, then Dynamo routes it to an available reasoning GPU. The reasoning GPU generates the reasoning chain tokens and stores the KV cache in a shared memory pool. When reasoning completes, the KV cache is transferred to the final-answer GPU for the output generation, freeing the reasoning GPU for the next query.

KV cache checkpointing to NVMe storage is essential for reasoning infrastructure. A 20,000-token KV cache at 5.1GB takes approximately 5-10 seconds to write to a modern Gen5 NVMe drive (5-10 GB/s sequential write). Periodic KV cache checkpointing (every 500-1,000 reasoning tokens) enables the reasoning GPU to be preempted for higher-priority queries or hardware maintenance without losing the reasoning chain progress. On spot GPU instances (which can be preempted with 2-minute notice), checkpointing every 200 reasoning tokens limits the work-at-risk to approximately $0.005 per failure (at $0.000319/sec H100 cost) versus $0.10+ if the full reasoning chain must restart.

The most expensive failure mode for reasoning serving: GPU OOM during long reasoning chains. A reasoning query that exhausts GPU HBM (80GB on H100 SXM5) mid-generation crashes the GPU process, losing all reasoning progress. Memory guards should be implemented at the application level: the serving framework (vLLM, SGLang) should monitor KV cache memory consumption during reasoning and preemptively terminate or offload the query when memory crosses a configurable threshold (typically 75-80% of total HBM). This prevents a single long reasoning chain from monopolizing GPU memory and blocking other queries.

07

The Future of Reasoning Infrastructure and Recommendations for 2026

Test-time compute scaling is not a passing trend - it is the direction of the entire LLM field. OpenAI's o-series, DeepSeek's R-series, Google's Gemini-Thinking, and Anthropic's reasoning-enhanced Claude all point toward models that internalize more computation at inference time. The infrastructure industry is responding with purpose-built reasoning serving architectures: B200's FP4 Tensor Cores, NVIDIA's Dynamo with reasoning-aware scheduling, and SGLang's reasoning-optimized serving branch. The GPU rental market on ClusterBid already shows operators segmenting inventory into 'reasoning-optimized' (high-HBM, high-bandwidth GPUs for decode) and 'standard' tiers.

For teams deploying reasoning models in production at mid-2026, the recommendations are: (1) budget for 5-10x more GPU capacity per query than you would for standard inference - the cost per reasoning query is $0.10-0.25 on H100, and this does not decrease as fast as training costs are decreasing; (2) implement separate GPU pools for prefill, reasoning decode, and final output to avoid head-of-line blocking where a long reasoning chain delays short queries; (3) set maximum reasoning token limits (32K-48K as a starting point) and implement KV cache checkpointing to protect against GPU failures and preemption.

The most important recommendation: profile your actual reasoning token distribution before committing to GPU capacity. The difference between a reasoning system averaging 5,000 tokens per query and one averaging 25,000 tokens is a 5x infrastructure cost difference. Run your model on 500 representative queries, measure the reasoning token distribution (P50, P90, P99), and size your GPU deployment for the P90 reasoning token length with a timeout at P99. This data-driven approach prevents both over-provisioning (buying GPUs for worst-case reasoning length) and under-provisioning (crashing during long reasoning chains). ClusterBid's flexible rental terms allow for capacity adjustments as your reasoning workload profile evolves and as model-level optimizations reduce per-query token consumption.

Filed under
Test-Time ComputeReasoning ModelsDeepSeek-R1KV CacheReasoning TokensInference CostGPU ConcurrencyChain-of-ThoughtScaling Laws