Open source · Runs in your browser
LLM VRAM Calculator
Paste a Hugging Face model id (or its URL) to load its layers, KV heads and parameter count from config.json, then pick a quant and a context length. You get weights, KV cache and overhead separately, and which GPU sizes fit.
Try: Qwen3.8 27B · Gemma 4 12B · gpt-oss-20b · Qwen3 30B-A3B
Or type the numbers below by hand.
—Weights (GiB)
—KV cache (GiB)
—Overhead (GiB)
—Estimated total (GiB)
Overhead is 0.5 GiB plus 10% of weights and cache. A GPU "fits" when at least 0.5 GiB stays free. Planning estimates, not measured peaks.
Which GPU sizes fit
| GPU memory | Fits? | Free after load | Longest context that fits |
|---|
Link to this result
This link opens the Space with the same model and settings. To put a VRAM badge on a model card, use the Markdown below. It needs no token and loads the model's own config.