§ GPU Recommender

How much GPU does your model need?

Search any model on HuggingFace, or start from the ones below. We size the weights and KV cache for your precision and context, then show every GPU configuration that fits.

FIG.1

Pick a model

Search every text-generation model on HuggingFace, or start from one below.

Popular models
billion parameters
FIG.2

Set the workload

Memory depends as much on how you run the model as on the model itself.

FIG.3

Memory footprint

Pick a model, or enter a parameter count, and the weights, KV cache and runtime overhead land here - with every GPU configuration that fits below.

Weights-
KV cache-
Runtime overhead-

Questions

How accurate are these numbers?

Weight memory is exact arithmetic: parameter count times bytes per parameter. The KV cache figure is exact when the model publishes its config, since it comes from the real layer and key/value head counts. What we approximate is the 15% we add for the CUDA context, the serving runtime, activation buffers and allocator fragmentation. Treat the total as a floor to plan against, not a guarantee. A production deployment with continuous batching and paged attention will land close to it; one with an unusual runtime may not.

Why is your number different from what I read elsewhere?

Most calculators multiply by the number of attention heads. Modern models use grouped-query attention, where key and value heads are far fewer than attention heads - Qwen2.5-7B has 28 attention heads but only 4 key/value heads. Using the wrong one overstates the KV cache by up to eight times. We use key/value heads, which is why our long-context figures are usually lower than a naive calculator's.

What happens with gated models like Llama?

HuggingFace still publishes the parameter count for gated repositories, so weight memory stays exact. The config file that carries layer and head counts returns a 401, so we estimate the attention geometry from parameter count using real reference models and label the KV cache line as approximate. Everything keeps working; you just get an estimate instead of an exact figure for that one component.

Why does fine-tuning need so much more memory than inference?

Full fine-tuning holds far more than the weights. Mixed-precision AdamW keeps 2 bytes of weights, 2 bytes of gradients, a 4-byte FP32 master copy and two 4-byte optimiser moments - roughly 16 bytes per parameter against inference's 2. A 7B model that serves on a single card needs about 112 GB of state before activations. LoRA avoids this by freezing the base weights and training a small fraction, which is why it fits where a full fine-tune does not.

Can I run a 70B model on one GPU?

At BF16 a 70B model is about 140 GB of weights, so no single card holds it. At INT4 it is roughly 35 GB, which fits on one 80 GB H100 or A100 with room for a modest KV cache, and comfortably on a 141 GB H200. The quality cost of INT4 is real but often acceptable for serving. The precision comparison in the tool shows the tradeoff directly.

How does context length change the answer?

Weights are fixed; the KV cache is not. It scales linearly with context length and with the number of concurrent requests, so one request at 128K tokens costs the same memory as sixteen at 8K. For high-concurrency serving the cache often exceeds the weights, which is where KV cache quantisation earns its place - dropping it to INT8 halves that component.

Which GPUs does it consider?

Only SKUs we actually source, from L4 and A10 up through H100, H200, B200, B300 and the GB-series racks. Every VRAM, bandwidth and interconnect figure comes from the NVIDIA datasheet, not from a vendor listing. If a configuration appears here, it is one we can quote.

What does Deploy now do?

It takes you to sign-in, where you pick the configuration and we place it with a provider that has the capacity. Pricing is deliberately not shown on this page: rates move with supply, term, region and commitment, and a single number here would be out of date or misleading more often than not. You get real numbers against your actual requirement.