SSM COMPUTE PROFILE VS TRANSFORMER
State space models replace the attention mechanism with a selective scan operation that processes tokens in a recurrent manner. Mamba-2 uses a structured SSM with a 1D convolutional front-end and a selective scan that scales linearly with sequence length, in contrast to the quadratic scaling of full attention. The key GPU implication: the selective scan is a memory-bandwidth-bound operation that achieves 35-45% of H100's peak HBM3 bandwidth of 3.35 TB/s on 256x256 state dimensions. In comparison, FlashAttention-3 achieves 55-65% of peak bandwidth on the same hardware. The scan kernel does not benefit from tensor core utilization (it's not a matrix multiply), which changes the GPU compute profile entirely.
For Mamba-2 2.8B on an H100, the selective scan takes 0.8 ms per forward pass at 8K sequence length versus 4.2 ms for FlashAttention-3 in a comparable Transformer 2.8B. At 128K context, the Mamba scan takes 12 ms while the Transformer attention takes 680 ms (FP8 with FlashAttention-3). This 50x advantage at long contexts is the core value proposition. However, the SSM's convolutional and projection layers still use standard matrix multiplies that benefit from tensor cores. Mamba-2 2.8B achieves 62% matmul utilization versus 71% for a comparable Transformer, because the attention-heavy Transformer spends more time on tensor core operations.
| Operation | Seq Len 8K | Seq Len 32K | Seq Len 128K | Seq Len 1M |
|---|---|---|---|---|
| Mamba Scan (2.8B) | 0.8 ms | 3.1 ms | 12 ms | 96 ms |
| FlashAttention-3 (2.8B) | 4.2 ms | 36 ms | 680 ms | 42,000 ms |
| Speedup Factor | 5.3x | 11.6x | 56.7x | 437x |
| Mamba Peak BW Util | 38% | 41% | 43% | 45% |
| FA3 Peak BW Util | 59% | 61% | 58% | 52% |
KV-CACHE-FREE INFERENCE: MEMORY ADVANTAGES
The most impactful GPU memory difference: SSMs maintain a fixed-size state (typically 16 * state_dim * d_model bytes per layer, roughly 0.5-1 MB per layer for a 2.8B model) that does not grow with sequence length. A Transformer of the same size requires KV cache of 2 * n_layers * d_model * seq_len * precision bytes. At 128K context with 2.8B parameters (32 layers, 4,096 hidden dim, FP16), the Transformer KV cache is 32 * 2 * 4096 * 128K * 2 = 64 GB. The Mamba-2 state across all layers at 128K is 32 * 0.5 MB = 16 MB. This 4,000x memory difference means SSMs can serve 128K-context prompts on a single H100 while the Transformer requires 4-8 H100s just for KV cache.
In practice, the memory savings enable higher batch sizes on the same GPU. On a single H100 with 80 GB, Mamba-2 2.8B serves batch size 128 at 128K context: 4.2 GB for weights + 2 GB for states = 6.2 GB total, leaving 73 GB for activations and overhead. The Transformer 2.8B with the same context hits 64 GB KV cache + 5.6 GB weights = 69.6 GB, leaving only 10 GB and allowing batch size at most 8. For throughput-oriented workloads, Mamba achieves 16x higher batch throughput per GPU at 128K context. The Jamba hybrid architecture combines 1 SSM layer per transformer layer, holding KV cache for transformer layers only, reducing memory by roughly 50% versus pure Transformer.
| Metric | Mamba-2 2.8B | Transformer 2.8B | Jamba 3B (Hybrid) |
|---|---|---|---|
| Weights (FP16) | 5.6 GB | 5.6 GB | 6.0 GB |
| KV Cache at 128K | 16 MB (state) | 64 GB | 32 GB |
| Max BS at 128K on H100 | 128+ | 8 | 16 |
| Max BS at 8K on H100 | 256+ | 64 | 96 |
| Tokens/sec BS=32, 128K | 12,400 | 280 | 640 |
| Cost/1M tok BS=32, 128K | $0.03 | $1.34 | $0.58 |
THROUGHPUT AND LATENCY ACROSS GPUS
Mamba-2's throughput advantage varies by sequence length and GPU type. On H100 SXM 80 GB with batch size 64 at 8K context, Mamba-2 2.8B achieves 18,200 output tokens/second, versus 9,450 for the equivalent Transformer 2.8B. The 1.9x advantage comes from the scan operation being faster than attention, plus the memory savings allowing larger batch sizes. At 128K context, the advantage grows to 44x: Mamba at 12,400 tok/s versus Transformer at 280 tok/s, because the Transformer is thrashing KV cache and operating at tiny batch sizes. On A100 (80 GB PCIe), Mamba-2 still achieves 8,200 tok/s at 128K context, while the Transformer drops to 45 tok/s.
Cost per 1M tokens on H100 at $2.50/hr shows the economic divergence. Mamba-2 2.8B at 8K context costs $0.02 per 1M tokens. The Transformer costs $0.04. At 128K context, Mamba costs $0.03 versus Transformer at $1.34 per 1M tokens. For long-context RAG workloads processing 50M tokens monthly, Mamba inference costs $1,500 versus $67,000 for the Transformer on the same hardware. Jamba hybrid sits in the middle: $0.08 at 8K and $0.58 at 128K, offering a pragmatic middle ground for teams that need partial compatibility with the Transformer ecosystem while capturing most of the memory benefit.
MAMBA-3 ON B200: BANDWIDTH VS COMPUTE TRADEOFFS
Mamba-3 scales to 7B parameters with an expanded state dimension of 512 (versus 256 in Mamba-2), increasing state size per layer to 1.2 MB but maintaining the fixed-size advantage. On B200 with 1.8 TB/s HBM3e bandwidth (compared to H100's 3.35 TB/s), the selective scan operations are bandwidth-bound. Mamba-3 achieves 40-48% of B200's lower bandwidth, yielding approximately 720 GB/s effective throughput versus H100's 1,340 GB/s. This means Mamba-3 on B200 runs at roughly 54% of H100's scan throughput per unit time, while Transformer workloads on B200 benefit from more tensor core compute capability.
The surprising finding: for SSM inference, H100 outperforms B200 despite B200 being a newer architecture, because B200's bandwidth-to-compute ratio is worse for memory-bound SSM operations. B200 achieves 94 TFLOPS per TB/s of bandwidth, H100 achieves 50 TFLOPS per TB/s. SSMs are memory-bandwidth-bound, making H100 the better GPU for pure SSM inference. The Jamba hybrid model partially addresses this because the transformer layers benefit from B200's tensor core improvements, making B200 viable for hybrid architectures. For teams running pure SSM models, H100 delivers 1.5-1.8x better cost-performance at $2.50/hr versus B200 at $3.80/hr.
