Calculator · Local AI & Hardware
How much VRAM do you need to run an LLM locally?
Short answer
Weights take parameters × bits per weight ÷ 8: about 5 GB for an 8B model at Q4_K_M, 20 GB for a 32B model. Add the KV cache, which grows with the context (1 GiB per 8,192 tokens on Llama 3.1 8B), and about 1 GB of margin. So 8 GB fits an 8B model with a short context, 16 GB fits 14B to 20B models, and 70B models need more than any single consumer card.

Prices, limits and features change often. We date every figure and link to its source: check the vendor's page before you buy or build. How we make money.
A language model running on your own computer needs memory for three things: its weights, the KV cache that holds the conversation so far, and a margin for the engine itself. The weights depend on the model's size and how heavily it is compressed (quantized). The KV cache depends on the model's architecture and on how much context you give it. That second part is the one most guides skip, and it is why a model that "fits" can still fail when you paste in a long document.
The calculator below uses the same formula we checked against 13 real model files (details in How we checked).
Estimated memory needed
7.3 GB
- Weights
- 5.0 GB
- KV cache
- 1.2 GB
- Margin
- 1.1 GB
- Smallest common memory size that fits8 GB
- NVIDIA GeForce cards with enough VRAM
Theoretical ceiling on an RTX 5090: about 358 tokens per second (memory bandwidth ÷ weight size). Real speed is lower.
- Weights = parameters × bits per weight ÷ 8, with the bits per weight that llama.cpp measured for each quantization.
- KV cache = 2 × layers that keep the full context × KV heads × head dimension × bytes per value × tokens. Sliding-window layers stop at their window.
- The 1 GiB margin is our assumption for compute buffers and the driver. It varies with the engine and settings.
- Values in GB (10⁹ bytes), the unit GPU makers use. Nothing you enter leaves your browser.
The formula in three lines
- Weights (bytes) = parameters × bits per weight ÷ 8. An 8-billion-parameter model at Q4_K_M (4.89 bits per weight, as measured by llama.cpp) needs about 5 GB.
- KV cache (bytes) = 2 × layers × KV heads × head dimension × bytes per value × tokens of context. The 2 is for keys and values. At the default 16-bit precision, each value takes 2 bytes.
- Margin: we add 1 GiB for compute buffers and the GPU driver. This is our assumption, not a measurement. The real figure depends on the engine, the batch size and the GPU.
The architecture numbers come from each model's config.json on Hugging Face (num_hidden_layers, num_key_value_heads, head_dim). The calculator has them built in for 12 popular open models; for any other model, choose "Other model" and copy them in.
How much memory popular models need
With an 8,192-token context, which is enough for a long chat or a few pages of text:
| Model | Parameters | Q4_K_M | Q8_0 | Fits in (Q4_K_M) |
|---|---|---|---|---|
| Qwen3 4B | 4 B | 4.7 | 6.6 | 8 GB |
| Qwen3 8B | 8.2 B | 7.3 | 11.0 | 8 GB |
| Llama 3.1 8B | 8 B | 7.1 | 10.7 | 8 GB |
| Gemma 3 12B | 12.2 B | 9.4 | 14.9 | 12 GB |
| Qwen3 14B | 14.8 B | 11.5 | 18.1 | 12 GB |
| gpt-oss 20B | 20.9 B, (MoE, 3.6 B active) | 15.1 MXFP4 (published) | 16 GB | |
| Mistral Small 3.2 24B | 24 B | 17.1 | 27.9 | 24 GB |
| Qwen3-Coder 30B-A3B | 30.5 B, (MoE, 3.3 B active) | 20.6 | 34.3 | 24 GB |
| Qwen3 32B | 32.8 B | 23.3 | 38.0 | 24 GB |
| Qwen3.6 35B-A3B | 36 B, (MoE, 3 B active) | 23.2 | 39.4 | 24 GB |
| Llama 3.3 70B | 70.6 B | 46.9 | 78.7 | 48 GB |
| gpt-oss 120B | 116.8 B, (MoE, 5.1 B active) | 66.7 MXFP4 (published) | 96 GB | |
Three patterns stand out:
- 8 GB cards are tight for 8B models. Qwen3 8B at Q4_K_M comes to about 7.3 GB with 8K of context, so it fits, but there is little room for a longer context or for anything else using the GPU.
- 16 GB covers most mid-size models at Q4_K_M, including Qwen3 14B (11.5 GB) and OpenAI's gpt-oss-20b (15.1 GB). OpenAI's own model card says gpt-oss-20b runs "within 16GB of memory", which matches our estimate.
- 70B models do not fit a single consumer card. Llama 3.3 70B needs about 43 GB just for its Q4_K_M weights. On a 24 GB or 32 GB GPU, part of the model has to run from system RAM, which is much slower.
Quantization: what Q4_K_M and Q8_0 mean
Quantization stores each weight with fewer bits. llama.cpp publishes the measured bits per weight for each format on Llama 3.1 8B:
| Format | Bits per weight | Llama 3.1 8B size |
|---|---|---|
| Q3_K_M | 4.00 | 3.74 GiB |
| Q4_K_M | 4.89 | 4.58 GiB |
| Q5_K_M | 5.70 | 5.33 GiB |
| Q6_K | 6.56 | 6.14 GiB |
| Q8_0 | 8.50 | 7.95 GiB |
| F16 | 16.00 | 14.96 GiB |
The bits per weight are above the nominal 4, 5 or 8 because each block of weights also stores scaling factors, and the "K" mixes keep some layers at higher precision. llama.cpp notes that quantization "may introduce some accuracy loss", which it measures with perplexity and KL divergence. How much that matters depends on the model and the task, so if a model is borderline, it is worth trying the next format up before buying more memory.
Some models are published already quantized. OpenAI's gpt-oss models ship with their mixture-of-experts weights in MXFP4, so the calculator uses the real file size for them (13.8 GB for gpt-oss-20b, 65.4 GB for gpt-oss-120b) and ignores the quantization setting.
Context length is the hidden cost
The KV cache grows in a straight line with the context. For Llama 3.1 8B at 16-bit precision it takes exactly 1 GiB at 8,192 tokens, 4 GiB at 32,768 and 16 GiB at its full 131,072 tokens. At full context, the cache is three and a half times the size of the model's own Q4_K_M weights.
Engines pick a default for you. Ollama's documentation says it uses a 4K context below 24 GiB of VRAM, 32K between 24 and 48 GiB, and 256K above that, and recommends at least 64,000 tokens for agents, web search and coding tools. If you raise it, you can see whether the model still fits with ollama ps: the PROCESSOR column shows "100% GPU" when it does.
Two settings shrink the cache:
- Quantized KV cache. Ollama's
OLLAMA_KV_CACHE_TYPEacceptsf16(the default),q8_0(about half the memory, which Ollama says "usually has no noticeable impact" on quality) andq4_0(about a quarter, with a loss that "may be more noticeable at higher context sizes"). It needs Flash Attention, which Ollama turns on automatically when the hardware supports it. - A shorter context. If your prompts are short, 4K or 8K is enough, and the memory you save can go to a better quantization.
Why some new models need far less cache
Many recent models do not keep the whole context in every layer. Some layers use a sliding window and only look at the last few hundred or thousand tokens. Others use linear attention, whose memory does not grow with the context at all. Calculators that treat every layer as full attention overestimate these models badly:
| Model | Layers with full context | Cache at 128K, counted correctly | If every layer kept the full context |
|---|---|---|---|
| gpt-oss-20b | 12 of 24 (the rest: 128-token window) | 3.2 GB | 6.4 GB |
| Gemma 3 12B | 8 of 48 (the rest: 1,024-token window) | 8.9 GB | 51.5 GB |
| Qwen3.6 35B-A3B (at 256K) | 10 of 40 (the rest: linear attention) | 5.4 GB | 21.5 GB |
These savings depend on the engine. llama.cpp keeps a reduced cache for sliding-window layers by default (its --swa-full option, which turns this off, defaults to false). If you use an engine that allocates the full cache for every layer, expect the larger figure.
Mixture-of-experts models: memory for all, speed of a few
Models such as Qwen3-Coder 30B-A3B or gpt-oss-20b are mixtures of experts. Only a few experts work on each token: 3.3 billion of Qwen3-Coder's 30.5 billion parameters, and 3.6 billion of gpt-oss-20b's 21 billion, according to their model cards. All the experts still have to be loaded, so memory follows the total size. Speed follows the active size, which is why these models feel much faster than a dense model of the same file size.
Which GPU: VRAM and bandwidth
NVIDIA's comparison table lists the memory of each GeForce card:
| Card | VRAM | Bandwidth |
|---|---|---|
| GeForce RTX 5090 | 32 GB | 1,792 GB/s |
| GeForce RTX 5080 | 16 GB | 960 GB/s |
| GeForce RTX 5070 Ti | 16 GB | 896 GB/s |
| GeForce RTX 5070 | 12 GB | 672 GB/s |
| GeForce RTX 5060 Ti | 16 / 8 GB | 448 GB/s |
| GeForce RTX 5060 | 8 GB | 448 GB/s |
| GeForce RTX 5050 | 8 GB | 320 GB/s |
| GeForce RTX 4090 | 24 GB | not listed |
| GeForce RTX 4080 SUPER | 16 GB | not listed |
| GeForce RTX 4080 | 16 GB | not listed |
| GeForce RTX 4070 Ti SUPER | 16 GB | not listed |
| GeForce RTX 4070 Ti | 12 GB | not listed |
| GeForce RTX 4070 SUPER | 12 GB | not listed |
| GeForce RTX 4070 | 12 GB | not listed |
| GeForce RTX 4060 Ti | 16 / 8 GB | not listed |
| GeForce RTX 4060 | 8 GB | not listed |
VRAM decides what fits. Memory bandwidth decides how fast it runs: to write each new token, the GPU reads every active weight once. That gives a hard ceiling of bandwidth ÷ weight size. For Qwen3 8B at Q4_K_M (5.0 GB), the ceiling is about 358 tokens per second on an RTX 5090 (1,792 GB/s) and about 89 on an RTX 5060 (448 GB/s). Real speeds are lower, because the GPU also reads the cache and does other work, but the ratio between cards holds. We have not measured real speeds ourselves: these are theoretical limits, not benchmarks.
On Apple Silicon Macs, the GPU uses the same unified memory as the rest of the system, so the relevant number is your Mac's total memory minus what macOS and your apps use. LM Studio recommends at least 16 GB on a Mac and says 8 GB machines should "stick to smaller models and modest context sizes". If you are choosing between apps, see Ollama vs LM Studio.
When the model does not fit
You have four options, roughly from best to worst:
- Use a smaller quantization (Q4_K_M instead of Q6_K), or a smaller model from the same family.
- Shorten the context or quantize the KV cache.
- Split the model between GPU and CPU. Ollama does this automatically when a model does not fit in VRAM, and
ollama psshows the split. It works, but every layer on the CPU slows down generation. - Add a GPU. Ollama spreads a model across several GPUs when it does not fit on one, according to its FAQ.
For a step-by-step setup, see how to run an LLM locally.
How we checked
We compared the weights formula against the real size of 13 GGUF files that unsloth publishes on Hugging Face, from Qwen3 4B to Llama 3.3 70B, at Q4_K_M and Q8_0. Every estimate fell within 3% of the real file:
| File | Real size | Our estimate | Difference |
|---|---|---|---|
| Qwen3-8B Q4_K_M | 5.03 GB | 5.01 GB | -0.3% |
| Qwen3-8B Q8_0 | 8.71 GB | 8.70 GB | -0.1% |
| Llama-3.1-8B Q4_K_M | 4.92 GB | 4.91 GB | -0.2% |
| Qwen3-14B Q4_K_M | 9.00 GB | 9.04 GB | +0.4% |
| Qwen3-32B Q4_K_M | 19.76 GB | 20.04 GB | +1.4% |
| Llama-3.3-70B Q4_K_M | 42.52 GB | 43.16 GB | +1.5% |
| Gemma-3-12B Q4_K_M | 7.30 GB | 7.46 GB | +2.1% |
| Mistral-Small-3.2-24B Q4_K_M | 14.33 GB | 14.69 GB | +2.5% |
Gemma 3 and Mistral Small come out slightly high because their parameter count includes a vision encoder that ships in a separate file. The KV cache formula is exact arithmetic on the published architecture; the test suite behind this page checks it against known values (Llama 3.1 8B: 16 GiB at 131,072 tokens). What we have not measured is the runtime margin, which we set at 1 GiB, and real-world speed.
What we checked
- Q4_K_M uses 4.8944 bits per weight and Llama 3.1 8B at Q4_K_M takes 4.58 GiB; Q8_0 uses 8.5008 bits per weight. (ggml-org / llama.cpp, )
- Qwen3-8B has 36 layers, 8 KV heads and a head dimension of 128 (config.json). (Qwen (Hugging Face), )
- Llama 3.1 8B has 32 layers, 8 KV heads and a head dimension of 128 (config.json). (unsloth (Hugging Face), )
- Qwen3-32B has 64 layers, 8 KV heads and a head dimension of 128 (config.json). (Qwen (Hugging Face), )
- gpt-oss-20b has 21B parameters with 3.6B active, alternates sliding-window (128 tokens) and full-attention layers, and runs within 16GB of memory with MXFP4 weights. (OpenAI (Hugging Face), )
- Gemma 3 12B uses a 1,024-token sliding window with one global layer in six (sliding_window_pattern 6). (unsloth (Hugging Face), )
- Qwen3.6-35B-A3B has 35B parameters with 3B activated and 10 full-attention layers out of 40. (Qwen (Hugging Face), )
- Qwen3-Coder-30B-A3B has 30.5B parameters with 3.3B activated. (Qwen (Hugging Face), )
- Real GGUF sizes used to validate the formula (13 files, Qwen3, Llama 3.1/3.3, Gemma 3, Mistral Small 3.2, gpt-oss). (unsloth (Hugging Face), )
- Ollama defaults to a 4K context below 24 GiB of VRAM, 32K from 24 to 48 GiB and 256K above, and recommends at least 64,000 tokens for agents, web search and coding tools. (Ollama, )
- OLLAMA_KV_CACHE_TYPE accepts f16 (default), q8_0 (about half the memory) and q4_0 (about a quarter), and requires Flash Attention. (Ollama, )
- llama.cpp's --swa-full option (full-size sliding-window cache) defaults to false. (ggml-org / llama.cpp, )
- RTX 5090: 32 GB GDDR7 and 1,792 GB/s; RTX 5060: 8 GB and 448 GB/s; RTX 4090: 24 GB. (NVIDIA, )
- LM Studio recommends 16GB+ RAM on Macs and says 8GB Macs should stick to smaller models and modest context sizes. (LM Studio, )
What may change
- New models change the presets: we add a model when its config.json and file sizes are published.
- Engines keep improving memory use (cache quantization, sliding-window caches), so real usage can be lower than our estimate.
- GPU line-ups change; the card table follows NVIDIA's comparison page.
Frequently asked questions
Can an 8 GB GPU run an 8B model?
Yes, at Q4_K_M with a short context. Qwen3 8B needs about 5.0 GB of weights plus 1.2 GB of KV cache at 8,192 tokens and our 1 GB margin, about 7.3 GB in total. Longer contexts or higher-precision formats will not fit.
Can I run a 70B model on a 24 GB or 32 GB GPU?
Not entirely on the GPU. Llama 3.3 70B needs about 43 GB just for its Q4_K_M weights. Engines such as Ollama can keep part of the model in system RAM, but every layer that runs on the CPU makes generation slower.
How much VRAM does gpt-oss-20b need?
Its weights are published in MXFP4 and take 13.8 GB. With a short context and our margin it comes to about 15 GB, consistent with OpenAI's statement that it runs within 16 GB of memory.
Does a longer context really need more VRAM?
Yes. The KV cache grows in a straight line with the context: on Llama 3.1 8B it is 1 GiB at 8,192 tokens and 16 GiB at 131,072. Models with sliding-window or linear attention grow much more slowly.
Is the 1 GB margin exact?
No. It is our assumption for compute buffers and the GPU driver. The real overhead depends on the engine, the batch size and the GPU, so leave some headroom if your estimate is close to your card's limit.
Sources
- llama.cpp: quantize (bits per weight and sizes by quantization type), ggml-org / llama.cpp. Accessed October 1, 2026.
- Qwen3-8B model card and config.json, Qwen (Hugging Face). Accessed October 1, 2026.
- Qwen3-32B model card and config.json, Qwen (Hugging Face). Accessed October 1, 2026.
- Llama 3.1 8B Instruct config.json (unsloth mirror of Meta's weights), unsloth (Hugging Face). Accessed October 1, 2026.
- gpt-oss-20b model card and config.json, OpenAI (Hugging Face). Accessed October 1, 2026.
- Gemma 3 12B config.json (unsloth mirror of Google's weights), unsloth (Hugging Face). Accessed October 1, 2026.
- Qwen3.6-35B-A3B model card and config.json, Qwen (Hugging Face). Accessed October 1, 2026.
- Qwen3-Coder-30B-A3B-Instruct model card, Qwen (Hugging Face). Accessed October 1, 2026.
- GGUF files and sizes for Qwen3, Llama 3.1/3.3, Gemma 3, Mistral Small 3.2 and gpt-oss, unsloth (Hugging Face). Accessed October 1, 2026.
- Ollama: context length defaults and settings, Ollama. Accessed October 1, 2026.
- Ollama FAQ: Flash Attention, K/V cache quantization, multiple GPUs, Ollama. Accessed October 1, 2026.
- llama.cpp server: command-line options (--swa-full, --flash-attn, cache types), ggml-org / llama.cpp. Accessed October 1, 2026.
- Compare GeForce graphics cards (memory configuration and bandwidth), NVIDIA. Accessed October 1, 2026.
- LM Studio: system requirements, LM Studio. Accessed October 1, 2026.
Spotted an error or an outdated price? Tell us and we will fix it.
Change history
- : First published, with a calculator for 12 open models validated against 13 real GGUF files.
Next review: .