All essays
MarketMARKET REPORTFEB 2026

Qwen 3.6 With 1M Native Context: VRAM Math, KV Cache Scaling, and What 1M-Token Inference Actually Costs

Qwen 3.6 launched with 1M-token native context. Real VRAM math and serving cost for 1M-token inference on available GPUs.

01

QWEN 3.6 AND 1M CONTEXT

Qwen 3.6 achieves 1M-token native context through YaRN-based RoPE scaling and grouped-query attention with streaming KV cache. A 7B parameter model fits 1M context in 72 GB VRAM at FP8. The 32B variant requires 256 GB+ for 1M context. This enables document-level understanding, full codebase analysis, and multi-hour conversation history in a single inference pass without RAG.

02

VRAM MATH FOR 1M CONTEXT

KV cache dominates VRAM at 1M context. A 7B model with GQA (8 KV heads) at FP8 requires 2 bytes x 8 heads x 128 dimensions x 1M tokens = 2.0 GB KV cache plus 7 GB weights = 9 GB at FP8. For 32B model with 16 KV heads: 4.1 GB KV cache plus 32 GB weights = 36 GB at FP8. Real-world overhead from padding and attention masks adds 20-40%. H100's 80 GB handles 7B with 32K-128K context comfortably; full 1M context requires multi-GPU.

ModelContextKV CacheWeightsTotal VRAMMin GPU
Qwen 3.6-7B128K0.25 GB7 GB9 GB1x L40S
Qwen 3.6-7B1M2.0 GB7 GB11 GB1x H100
Qwen 3.6-32B128K0.5 GB32 GB39 GB1x H100
Qwen 3.6-32B1M4.1 GB32 GB43 GB1x H200
Qwen 3.6-72B1M9.2 GB72 GB98 GB4x H100
03

KV CACHE SCALING

KV cache grows linearly with context length but super-linearly in practice due to increased padding. FlashAttention-3 reduces memory from O(n^2) to O(n) but attention computation still scales O(n^2). For 1M tokens, attention computation requires 2.8 PFLOPS versus 9 GFLOPS for 1K tokens-a 311x increase. Prefill latency reaches 30-60 seconds for 1M tokens on H100, making the prefill-decode disaggregation critical for interactive use.

04

GPU REQUIREMENTS

Qwen 3.6-7B with 1M context runs on single H100 with FlashAttention-3 and sliding window attention. Qwen 3.6-32B requires H200 or 2x H100 with tensor parallelism. Qwen 3.6-72B needs 4-8 H100s or 2-4 B200s. The 72B variant at 1M context costs $4.50-8.00/hr in GPU rental depending on configuration and provider.

05

COST MODELING

A single 1M-token query on Qwen 3.6-72B costs $0.08-0.15 in compute (prefill + decode on 4x H100 for 60-90 seconds). At 10,000 queries/month, total GPU cost is $8,000-15,000/month. KV cache offloading to CPU memory reduces cost by 40-60% for latency-tolerant workloads. Streaming and speculative decoding reduce per-query cost by 30-50% for multi-turn conversations.

06

OPTIMIZATIONS FOR LONG CONTEXT

KV cache quantization to INT4 reduces memory by 50% with 0.5-1% accuracy loss. Prefix caching avoids recomputing shared context across queries. Streaming LLM with attention sinks maintains coherent generation beyond 1M tokens while limiting KV cache to 256K sliding window. SGLang's RadixAttention provides 30-50% cache reuse for multi-turn conversations. For production 1M-context serving, disaggregated prefill-decode on Dynamo delivers 3-5x throughput versus monolithic vLLM.

Filed under
Qwen 3.61M ContextKV CacheLong ContextInference CostVRAM MathSelf-Host 1M