Why FA3 Extracts More from Hopper and Blackwell Than FA2 Ever Could
FlashAttention-3 shipped in 2024, targeting NVIDIA Hopper (SM90a), and became the default attention kernel in vLLM and SGLang by mid-2025. Flash Attention 3 GPU performance improvements are real and substantial - but they are not uniform across GPU configurations. That is the part no buyer's guide actually explains.
FA3 exploits three hardware capabilities that FA2 could not reach. Warp specialization on Hopper SM90a lets producer warps pipeline HBM fetches while consumer warps execute GEMM simultaneously. The Tensor Memory Accelerator (TMA) enables bulk async copies from global to shared memory with zero warp occupancy overhead - effectively hiding memory latency behind compute. And FP8 MMA instructions allow the Q, K, V projection matrices to execute at half the memory footprint with roughly 2x the GEMM throughput. Together these push Hopper SM utilization from around 35% (FA2 baseline) to north of 75% at long context lengths.
The hardware dependency matters for buyers. The FA3 feature set is present on both Hopper and Blackwell, but implementation depth varies. Hopper (H100, H200) runs the full FA3 stack on SM90a. Blackwell (B200) extends this with a 4th-gen Transformer Engine that supports deeper hardware FP8 accumulation, enabling a Blackwell-specific FA3 kernel path that compounds the gains. This is not the same workload running on different hardware - it is a qualitatively different execution pipeline.
H100 SXM vs PCIe: The 67% Bandwidth Gap Becomes a 40% Throughput Gap at 128K Tokens
The thing nobody tells you about H100 SXM vs PCIe for FA3 workloads is that the gap is larger than the raw TFLOP comparison implies. H100 SXM5 delivers 3.35 TB/s of HBM3 bandwidth. H100 PCIe delivers 2.0 TB/s. That 67% bandwidth advantage for SXM translates differently depending on context length.
At short context (4K tokens), attention is still partially compute-bound and the bandwidth gap compresses to a 15-20% throughput difference. At 32K tokens, the memory wall starts dominating. At 128K tokens - where FA3 really matters - attention becomes almost entirely memory-bandwidth-bound. FA3 on H100 SXM delivers roughly 2.7x the throughput of FA2 at 128K context. FA3 on H100 PCIe lands around 2.3x. Both are meaningful gains. But SXM compounds.
The practical math: H100 SXM spot rates run roughly 30-40% above equivalent PCIe configurations. At 128K context, you pay 35% more per GPU-hour and get 17% more throughput. That makes SXM the right call for any team running continuous batching with context windows above 32K, because the bandwidth covers the long-tail requests that would otherwise throttle your batch utilization. For development workloads or short-context fine-tuning, H100 PCIe is fine.
B200 FP8 + FA3: Two Kernel Optimizations That Compound Rather Than Add
B200 changes the calculus. The bandwidth jump to 8.0 TB/s HBM3e (versus 3.35 TB/s on H100 SXM) matters, but the more important factor for FA3 is the 4th-gen Transformer Engine's native FP8 support. When running FA3 in FP8 mode on B200, the Q, K, V projections execute at FP8 precision with hardware-accelerated accumulation - doubling effective GEMM throughput and reducing KV cache memory footprint by roughly 50% versus FP16.
These two optimizations compound rather than add. FA3's async pipeline on B200 can schedule FP8 GEMM and softmax normalization concurrently in a way H100 cannot, because B200's hardware FP8 path runs directly through the Transformer Engine rather than going through the general FP8 emulation path. SM utilization on B200 at 128K context sits around 85-90% versus H100 SXM at roughly 75%. That 15-point utilization gap, across higher clock frequencies and 2x the SM count, produces a substantial effective throughput difference.
For FA3 H100 B200 throughput comparison at 128K context: B200 with FA3 FP8 delivers roughly 5x the attention throughput of FA2 on H100 SXM. If your primary workload is long-context inference - multi-document retrieval, agentic pipelines with large system prompts, code generation over large codebases - B200 SXM with FA3 FP8 is not a marginal upgrade. It changes what is economically viable to serve.
FA3 Throughput Multipliers at 4K, 32K, and 128K Context Lengths
These figures reflect approximate attention throughput multipliers versus FA2 baseline on H100 SXM, drawn from production deployments and published benchmarks. Exact gains vary with batch size, model architecture, and tensor parallelism degree - treat them as orientation for GPU selection, not SLA commitments.
The pattern is unambiguous: FA3 gains grow with context length, and B200's FP8 advantage compounds that growth. At 4K context, the configurations are close enough that pricing and availability dominate the decision. At 128K context, the gap between H100 PCIe and B200 FP8 is roughly 2.5x in FA3 throughput - which at current spot rates means B200 has a lower cost per inference despite the higher per-GPU-hour price.
| Context Length | H100 SXM (FA3 FP16) | B200 SXM (FA3 FP8) |
|---|---|---|
| 4K tokens | ~1.6x vs FA2 | ~2.1x vs FA2 |
| 32K tokens | ~2.2x vs FA2 | ~3.6x vs FA2 |
| 128K tokens | ~2.7x vs FA2 | ~5.1x vs FA2 |
H200: The Long-Context Sweet Spot Most Teams Overlook
H200 is consistently underweighted in these comparisons. Its 4.8 TB/s HBM3e bandwidth sits between H100 SXM (3.35 TB/s) and B200 (8.0 TB/s), and since FA3 is IO-bound at long contexts, bandwidth translates almost linearly to throughput. H200 delivers roughly 3.0x FA2 throughput at 128K context in FP16 mode - meaningfully better than H100 SXM and much cheaper than B200.
H200 runs the full Hopper FA3 stack (SM90a), including FP8 via the 3rd-gen Transformer Engine. The FP8 FA3 path on H200 delivers around 3.8x versus FA2 at 128K context. That is short of B200's 5.1x, but the gap reflects hardware FP8 accumulation depth, not a software limitation. H200 FP8 FA3 is a genuine and accessible configuration for 2026.
H200 spot pricing in mid-2026 runs roughly $2.80-3.16/GPU/hr on the open market. B200 SXM ranges from $3.00 to $6.03/GPU/hr depending on provider and commit term. For teams that need 32K-64K context capability and want to avoid B200 contract risk, H200 is often the right call. There is more H200 inventory in the market right now, and the FA3 gains on H200 hold up well through 64K context.
Flash Attention 3 GPU Selection Guide: Which Tier Is Worth Paying For
Here is the decision tree. P90 context under 16K tokens: H100 SXM is the call. FA3 gains are real but you do not need B200 memory bandwidth at this range. H100 SXM over PCIe for any multi-GPU serving setup because the NVLink topology matters for tensor parallelism. H100 PCIe makes sense only if you are single-GPU or budget-constrained and your context is genuinely short.
P90 context in the 32K-64K range: H200 or H100 SXM depending on your price sensitivity and whether your provider has H200 available. At this range, H200 delivers 15-20% better throughput per dollar than H100 SXM for FA3 FP16 workloads. If you are running mixed-length batches with occasional 128K requests, H200 absorbs those long-tail sequences without collapsing batch throughput.
P90 context at 128K or longer: B200 SXM. The math is not close. FA3 FP8 on B200 SXM delivers roughly 5x the throughput of FA2 on H100 SXM at 128K context, while B200 SXM costs roughly 2-2.5x the hourly rate. The cost-per-inference comparison favors B200 at this context length even accounting for the premium. One important caveat: B200 PCIe exists and costs less. We would still pick B200 SXM at the 25-30% form-factor premium, because at 128K context you are bandwidth-limited and trading memory bandwidth for hardware savings is the wrong tradeoff at this regime.
Finding B200 SXM Inventory Faster Than Direct Channels
If you have worked through this analysis and landed on B200 SXM, the immediate problem is supply. Direct provider channels for B200 SXM have 36-52 week lead times through most of 2026. The hyperscalers have capacity but their Blackwell availability is dominated by reserved instances with multi-year commit requirements.
The spot market is more interesting. B200 SXM capacity is available from neoclouds and secondary market providers, but fragmented - no single provider has it consistently and in volume. This is exactly the gap a broker model closes. ClusterBid sources B200 SXM inventory across dozens of providers and surfaces available capacity faster than direct channels can, without requiring you to work through each provider's sales process individually.
If your FA3 analysis points to B200 SXM for a specific workload, the bottleneck is usually not budget - it is finding supply at the moment you need it. Browse available GPU inventory at ClusterBid to see what is in the market right now.
