13k.euES
Menu

Calculator · Local AI & Hardware

How much VRAM do you need to run an LLM locally?

By 13k.eu editorsUpdated and checked 7 min read

Leer en español

Short answer

Weights take parameters × bits per weight ÷ 8: about 5 GB for an 8B model at Q4_K_M, 20 GB for a 32B model. Add the KV cache, which grows with the context (1 GiB per 8,192 tokens on Llama 3.1 8B), and about 1 GB of margin. So 8 GB fits an 8B model with a short context, 16 GB fits 14B to 20B models, and 70B models need more than any single consumer card.

Exploded view of a graphics card: fans, heat pipes, heatsink fins and the circuit board with the GPU glowing orange

Prices, limits and features change often. We date every figure and link to its source: check the vendor's page before you buy or build. How we make money.

A language model running on your own computer needs memory for three things: its weights, the KV cache that holds the conversation so far, and a margin for the engine itself. The weights depend on the model's size and how heavily it is compressed (quantized). The KV cache depends on the model's architecture and on how much context you give it. That second part is the one most guides skip, and it is why a model that "fits" can still fail when you paste in a long document.

The calculator below uses the same formula we checked against 13 real model files (details in How we checked).

Your model and settings (example values)

Q4_K_M is the usual default for local use; Q8_0 is close to the original quality and almost twice the size.

This model accepts up to 40,960 tokens. Ollama's default is 4K below 24 GiB of VRAM.

Estimated memory needed

7.3 GB

Weights
5.0 GB
KV cache
1.2 GB
Margin
1.1 GB
  • Smallest common memory size that fits8 GB
  • NVIDIA GeForce cards with enough VRAMRTX 5090 (32 GB) · RTX 5080 (16 GB) · RTX 5070 Ti (16 GB) · RTX 5070 (12 GB) · RTX 5060 Ti (16/8 GB) · RTX 5060 (8 GB) · RTX 5050 (8 GB) · RTX 4090 (24 GB) · RTX 4080 SUPER (16 GB) · RTX 4080 (16 GB) · RTX 4070 Ti SUPER (16 GB) · RTX 4070 Ti (12 GB) · RTX 4070 SUPER (12 GB) · RTX 4070 (12 GB) · RTX 4060 Ti (16/8 GB) · RTX 4060 (8 GB) · RTX 3090 Ti (24 GB) · RTX 3090 (24 GB) · RTX 3060 (12/8 GB)

Theoretical ceiling on an RTX 5090: about 358 tokens per second (memory bandwidth ÷ weight size). Real speed is lower.

  • Weights = parameters × bits per weight ÷ 8, with the bits per weight that llama.cpp measured for each quantization.
  • KV cache = 2 × layers that keep the full context × KV heads × head dimension × bytes per value × tokens. Sliding-window layers stop at their window.
  • The 1 GiB margin is our assumption for compute buffers and the driver. It varies with the engine and settings.
  • Values in GB (10⁹ bytes), the unit GPU makers use. Nothing you enter leaves your browser.

The formula in three lines

  • Weights (bytes) = parameters × bits per weight ÷ 8. An 8-billion-parameter model at Q4_K_M (4.89 bits per weight, as measured by llama.cpp) needs about 5 GB.
  • KV cache (bytes) = 2 × layers × KV heads × head dimension × bytes per value × tokens of context. The 2 is for keys and values. At the default 16-bit precision, each value takes 2 bytes.
  • Margin: we add 1 GiB for compute buffers and the GPU driver. This is our assumption, not a measurement. The real figure depends on the engine, the batch size and the GPU.

The architecture numbers come from each model's config.json on Hugging Face (num_hidden_layers, num_key_value_heads, head_dim). The calculator has them built in for 12 popular open models; for any other model, choose "Other model" and copy them in.

With an 8,192-token context, which is enough for a long chat or a few pages of text:

Estimated memory at 8,192 tokens of context (weights + f16 KV cache + 1 GiB margin), in GB
ModelParametersQ4_K_MQ8_0Fits in (Q4_K_M)
Qwen3 4B4 B4.76.68 GB
Qwen3 8B8.2 B7.311.08 GB
Llama 3.1 8B8 B7.110.78 GB
Gemma 3 12B12.2 B9.414.912 GB
Qwen3 14B14.8 B11.518.112 GB
gpt-oss 20B20.9 B, (MoE, 3.6 B active)15.1 MXFP4 (published)16 GB
Mistral Small 3.2 24B24 B17.127.924 GB
Qwen3-Coder 30B-A3B30.5 B, (MoE, 3.3 B active)20.634.324 GB
Qwen3 32B32.8 B23.338.024 GB
Qwen3.6 35B-A3B36 B, (MoE, 3 B active)23.239.424 GB
Llama 3.3 70B70.6 B46.978.748 GB
gpt-oss 120B116.8 B, (MoE, 5.1 B active)66.7 MXFP4 (published)96 GB

Three patterns stand out:

  1. 8 GB cards are tight for 8B models. Qwen3 8B at Q4_K_M comes to about 7.3 GB with 8K of context, so it fits, but there is little room for a longer context or for anything else using the GPU.
  2. 16 GB covers most mid-size models at Q4_K_M, including Qwen3 14B (11.5 GB) and OpenAI's gpt-oss-20b (15.1 GB). OpenAI's own model card says gpt-oss-20b runs "within 16GB of memory", which matches our estimate.
  3. 70B models do not fit a single consumer card. Llama 3.3 70B needs about 43 GB just for its Q4_K_M weights. On a 24 GB or 32 GB GPU, part of the model has to run from system RAM, which is much slower.

Quantization: what Q4_K_M and Q8_0 mean

Quantization stores each weight with fewer bits. llama.cpp publishes the measured bits per weight for each format on Llama 3.1 8B:

Format Bits per weight Llama 3.1 8B size
Q3_K_M 4.00 3.74 GiB
Q4_K_M 4.89 4.58 GiB
Q5_K_M 5.70 5.33 GiB
Q6_K 6.56 6.14 GiB
Q8_0 8.50 7.95 GiB
F16 16.00 14.96 GiB

The bits per weight are above the nominal 4, 5 or 8 because each block of weights also stores scaling factors, and the "K" mixes keep some layers at higher precision. llama.cpp notes that quantization "may introduce some accuracy loss", which it measures with perplexity and KL divergence. How much that matters depends on the model and the task, so if a model is borderline, it is worth trying the next format up before buying more memory.

Some models are published already quantized. OpenAI's gpt-oss models ship with their mixture-of-experts weights in MXFP4, so the calculator uses the real file size for them (13.8 GB for gpt-oss-20b, 65.4 GB for gpt-oss-120b) and ignores the quantization setting.

Context length is the hidden cost

The KV cache grows in a straight line with the context. For Llama 3.1 8B at 16-bit precision it takes exactly 1 GiB at 8,192 tokens, 4 GiB at 32,768 and 16 GiB at its full 131,072 tokens. At full context, the cache is three and a half times the size of the model's own Q4_K_M weights.

Engines pick a default for you. Ollama's documentation says it uses a 4K context below 24 GiB of VRAM, 32K between 24 and 48 GiB, and 256K above that, and recommends at least 64,000 tokens for agents, web search and coding tools. If you raise it, you can see whether the model still fits with ollama ps: the PROCESSOR column shows "100% GPU" when it does.

Two settings shrink the cache:

  • Quantized KV cache. Ollama's OLLAMA_KV_CACHE_TYPE accepts f16 (the default), q8_0 (about half the memory, which Ollama says "usually has no noticeable impact" on quality) and q4_0 (about a quarter, with a loss that "may be more noticeable at higher context sizes"). It needs Flash Attention, which Ollama turns on automatically when the hardware supports it.
  • A shorter context. If your prompts are short, 4K or 8K is enough, and the memory you save can go to a better quantization.

Why some new models need far less cache

Many recent models do not keep the whole context in every layer. Some layers use a sliding window and only look at the last few hundred or thousand tokens. Others use linear attention, whose memory does not grow with the context at all. Calculators that treat every layer as full attention overestimate these models badly:

Model Layers with full context Cache at 128K, counted correctly If every layer kept the full context
gpt-oss-20b 12 of 24 (the rest: 128-token window) 3.2 GB 6.4 GB
Gemma 3 12B 8 of 48 (the rest: 1,024-token window) 8.9 GB 51.5 GB
Qwen3.6 35B-A3B (at 256K) 10 of 40 (the rest: linear attention) 5.4 GB 21.5 GB

These savings depend on the engine. llama.cpp keeps a reduced cache for sliding-window layers by default (its --swa-full option, which turns this off, defaults to false). If you use an engine that allocates the full cache for every layer, expect the larger figure.

Mixture-of-experts models: memory for all, speed of a few

Models such as Qwen3-Coder 30B-A3B or gpt-oss-20b are mixtures of experts. Only a few experts work on each token: 3.3 billion of Qwen3-Coder's 30.5 billion parameters, and 3.6 billion of gpt-oss-20b's 21 billion, according to their model cards. All the experts still have to be loaded, so memory follows the total size. Speed follows the active size, which is why these models feel much faster than a dense model of the same file size.

Which GPU: VRAM and bandwidth

NVIDIA's comparison table lists the memory of each GeForce card:

Memory of NVIDIA GeForce cards (manufacturer's comparison table)
CardVRAMBandwidth
GeForce RTX 509032 GB1,792 GB/s
GeForce RTX 508016 GB960 GB/s
GeForce RTX 5070 Ti16 GB896 GB/s
GeForce RTX 507012 GB672 GB/s
GeForce RTX 5060 Ti16 / 8 GB448 GB/s
GeForce RTX 50608 GB448 GB/s
GeForce RTX 50508 GB320 GB/s
GeForce RTX 409024 GBnot listed
GeForce RTX 4080 SUPER16 GBnot listed
GeForce RTX 408016 GBnot listed
GeForce RTX 4070 Ti SUPER16 GBnot listed
GeForce RTX 4070 Ti12 GBnot listed
GeForce RTX 4070 SUPER12 GBnot listed
GeForce RTX 407012 GBnot listed
GeForce RTX 4060 Ti16 / 8 GBnot listed
GeForce RTX 40608 GBnot listed

VRAM decides what fits. Memory bandwidth decides how fast it runs: to write each new token, the GPU reads every active weight once. That gives a hard ceiling of bandwidth ÷ weight size. For Qwen3 8B at Q4_K_M (5.0 GB), the ceiling is about 358 tokens per second on an RTX 5090 (1,792 GB/s) and about 89 on an RTX 5060 (448 GB/s). Real speeds are lower, because the GPU also reads the cache and does other work, but the ratio between cards holds. We have not measured real speeds ourselves: these are theoretical limits, not benchmarks.

On Apple Silicon Macs, the GPU uses the same unified memory as the rest of the system, so the relevant number is your Mac's total memory minus what macOS and your apps use. LM Studio recommends at least 16 GB on a Mac and says 8 GB machines should "stick to smaller models and modest context sizes". If you are choosing between apps, see Ollama vs LM Studio.

When the model does not fit

You have four options, roughly from best to worst:

  1. Use a smaller quantization (Q4_K_M instead of Q6_K), or a smaller model from the same family.
  2. Shorten the context or quantize the KV cache.
  3. Split the model between GPU and CPU. Ollama does this automatically when a model does not fit in VRAM, and ollama ps shows the split. It works, but every layer on the CPU slows down generation.
  4. Add a GPU. Ollama spreads a model across several GPUs when it does not fit on one, according to its FAQ.

For a step-by-step setup, see how to run an LLM locally.

How we checked

We compared the weights formula against the real size of 13 GGUF files that unsloth publishes on Hugging Face, from Qwen3 4B to Llama 3.3 70B, at Q4_K_M and Q8_0. Every estimate fell within 3% of the real file:

File Real size Our estimate Difference
Qwen3-8B Q4_K_M 5.03 GB 5.01 GB -0.3%
Qwen3-8B Q8_0 8.71 GB 8.70 GB -0.1%
Llama-3.1-8B Q4_K_M 4.92 GB 4.91 GB -0.2%
Qwen3-14B Q4_K_M 9.00 GB 9.04 GB +0.4%
Qwen3-32B Q4_K_M 19.76 GB 20.04 GB +1.4%
Llama-3.3-70B Q4_K_M 42.52 GB 43.16 GB +1.5%
Gemma-3-12B Q4_K_M 7.30 GB 7.46 GB +2.1%
Mistral-Small-3.2-24B Q4_K_M 14.33 GB 14.69 GB +2.5%

Gemma 3 and Mistral Small come out slightly high because their parameter count includes a vision encoder that ships in a separate file. The KV cache formula is exact arithmetic on the published architecture; the test suite behind this page checks it against known values (Llama 3.1 8B: 16 GiB at 131,072 tokens). What we have not measured is the runtime margin, which we set at 1 GiB, and real-world speed.

What we checked

  • Q4_K_M uses 4.8944 bits per weight and Llama 3.1 8B at Q4_K_M takes 4.58 GiB; Q8_0 uses 8.5008 bits per weight. (ggml-org / llama.cpp, )
  • Qwen3-8B has 36 layers, 8 KV heads and a head dimension of 128 (config.json). (Qwen (Hugging Face), )
  • Llama 3.1 8B has 32 layers, 8 KV heads and a head dimension of 128 (config.json). (unsloth (Hugging Face), )
  • Qwen3-32B has 64 layers, 8 KV heads and a head dimension of 128 (config.json). (Qwen (Hugging Face), )
  • gpt-oss-20b has 21B parameters with 3.6B active, alternates sliding-window (128 tokens) and full-attention layers, and runs within 16GB of memory with MXFP4 weights. (OpenAI (Hugging Face), )
  • Gemma 3 12B uses a 1,024-token sliding window with one global layer in six (sliding_window_pattern 6). (unsloth (Hugging Face), )
  • Qwen3.6-35B-A3B has 35B parameters with 3B activated and 10 full-attention layers out of 40. (Qwen (Hugging Face), )
  • Qwen3-Coder-30B-A3B has 30.5B parameters with 3.3B activated. (Qwen (Hugging Face), )
  • Real GGUF sizes used to validate the formula (13 files, Qwen3, Llama 3.1/3.3, Gemma 3, Mistral Small 3.2, gpt-oss). (unsloth (Hugging Face), )
  • Ollama defaults to a 4K context below 24 GiB of VRAM, 32K from 24 to 48 GiB and 256K above, and recommends at least 64,000 tokens for agents, web search and coding tools. (Ollama, )
  • OLLAMA_KV_CACHE_TYPE accepts f16 (default), q8_0 (about half the memory) and q4_0 (about a quarter), and requires Flash Attention. (Ollama, )
  • llama.cpp's --swa-full option (full-size sliding-window cache) defaults to false. (ggml-org / llama.cpp, )
  • RTX 5090: 32 GB GDDR7 and 1,792 GB/s; RTX 5060: 8 GB and 448 GB/s; RTX 4090: 24 GB. (NVIDIA, )
  • LM Studio recommends 16GB+ RAM on Macs and says 8GB Macs should stick to smaller models and modest context sizes. (LM Studio, )

What may change

  • New models change the presets: we add a model when its config.json and file sizes are published.
  • Engines keep improving memory use (cache quantization, sliding-window caches), so real usage can be lower than our estimate.
  • GPU line-ups change; the card table follows NVIDIA's comparison page.

Frequently asked questions

Can an 8 GB GPU run an 8B model?

Yes, at Q4_K_M with a short context. Qwen3 8B needs about 5.0 GB of weights plus 1.2 GB of KV cache at 8,192 tokens and our 1 GB margin, about 7.3 GB in total. Longer contexts or higher-precision formats will not fit.

Can I run a 70B model on a 24 GB or 32 GB GPU?

Not entirely on the GPU. Llama 3.3 70B needs about 43 GB just for its Q4_K_M weights. Engines such as Ollama can keep part of the model in system RAM, but every layer that runs on the CPU makes generation slower.

How much VRAM does gpt-oss-20b need?

Its weights are published in MXFP4 and take 13.8 GB. With a short context and our margin it comes to about 15 GB, consistent with OpenAI's statement that it runs within 16 GB of memory.

Does a longer context really need more VRAM?

Yes. The KV cache grows in a straight line with the context: on Llama 3.1 8B it is 1 GiB at 8,192 tokens and 16 GiB at 131,072. Models with sliding-window or linear attention grow much more slowly.

Is the 1 GB margin exact?

No. It is our assumption for compute buffers and the GPU driver. The real overhead depends on the engine, the batch size and the GPU, so leave some headroom if your estimate is close to your card's limit.

Sources

  1. llama.cpp: quantize (bits per weight and sizes by quantization type), ggml-org / llama.cpp. Accessed October 1, 2026.
  2. Qwen3-8B model card and config.json, Qwen (Hugging Face). Accessed October 1, 2026.
  3. Qwen3-32B model card and config.json, Qwen (Hugging Face). Accessed October 1, 2026.
  4. Llama 3.1 8B Instruct config.json (unsloth mirror of Meta's weights), unsloth (Hugging Face). Accessed October 1, 2026.
  5. gpt-oss-20b model card and config.json, OpenAI (Hugging Face). Accessed October 1, 2026.
  6. Gemma 3 12B config.json (unsloth mirror of Google's weights), unsloth (Hugging Face). Accessed October 1, 2026.
  7. Qwen3.6-35B-A3B model card and config.json, Qwen (Hugging Face). Accessed October 1, 2026.
  8. Qwen3-Coder-30B-A3B-Instruct model card, Qwen (Hugging Face). Accessed October 1, 2026.
  9. GGUF files and sizes for Qwen3, Llama 3.1/3.3, Gemma 3, Mistral Small 3.2 and gpt-oss, unsloth (Hugging Face). Accessed October 1, 2026.
  10. Ollama: context length defaults and settings, Ollama. Accessed October 1, 2026.
  11. Ollama FAQ: Flash Attention, K/V cache quantization, multiple GPUs, Ollama. Accessed October 1, 2026.
  12. llama.cpp server: command-line options (--swa-full, --flash-attn, cache types), ggml-org / llama.cpp. Accessed October 1, 2026.
  13. Compare GeForce graphics cards (memory configuration and bandwidth), NVIDIA. Accessed October 1, 2026.
  14. LM Studio: system requirements, LM Studio. Accessed October 1, 2026.

Spotted an error or an outdated price? Tell us and we will fix it.

Change history

  • : First published, with a calculator for 12 open models validated against 13 real GGUF files.

Next review: .

Part of our Local AI & Hardware guide.