Open source · Runs in your browser

LLM VRAM Calculator

Paste a Hugging Face model id (or its URL) to load its layers, KV heads and parameter count from config.json, then pick a quant and a context length. You get weights, KV cache and overhead separately, and which GPU sizes fit.

Try: Qwen3.8 27B · Gemma 4 12B · gpt-oss-20b · Qwen3 30B-A3B

Or type the numbers below by hand.

—Weights (GiB)
—KV cache (GiB)
—Overhead (GiB)
—Estimated total (GiB)

Overhead is 0.5 GiB plus 10% of weights and cache. A GPU "fits" when at least 0.5 GiB stays free. Planning estimates, not measured peaks.

Which GPU sizes fit

GPU memoryFits?Free after loadLongest context that fits
Open in the full calculator (real GGUF file sizes, MoE offload, llama-server command) →

Link to this result

This link opens the Space with the same model and settings. To put a VRAM badge on a model card, use the Markdown below. It needs no token and loads the model's own config.