LLM Hardware Tool

VRAM
Calculator.

Select a language model, quantization, and context length to calculate exactly how much VRAM your GPU needs.

Required VRAM

5.21GB

Breakdown

Model weights4.21 GB
KV Cache (ctx 4,096)0.50 GB
Runtime overhead0.50 GB
Total estimated5.21 GB

Recommended GPU

RTX 5060 8GB

8GB VRAM — Runs locally

shopping_cart View GPU on Amazon
check_circle

Minimum recommended: RTX 4060

Estimation based on Ollama / llama.cpp. Results may vary by framework.

article

New to local LLMs?

Read our full guide on how to run models on your PC with Ollama.

Model

DeepSeek R1 1.5Bctx 131K
1.78Bparams
DeepSeek R1 7Bctx 131K
7.6Bparams
DeepSeek R1 8Bctx 131K
8.03Bparams
DeepSeek R1 14Bctx 131K
14.7Bparams
DeepSeek R1 32Bctx 131K
32.5Bparams
DeepSeek R1 70Bctx 131K
70.6Bparams
Falcon 40Bctx 2K
40Bparams
Falcon 180Bctx 2K
180Bparams
Gemma 1 2Bctx 8K
2Bparams
Gemma 2 2Bctx 8K
2.6Bparams
Gemma 1 7Bctx 8K
7Bparams
Gemma 2 9Bctx 8K
9.24Bparams
Gemma 2 27Bctx 8K
27.2Bparams
Llama 3.2 1Bctx 131K
1.24Bparams
Llama 3.2 3Bctx 131K
3.21Bparams
Llama 1 7Bctx 2K
7Bparams
Llama 2 7Bctx 4K
7Bparams
Llama 3 8Bctx 8K
8Bparams
Llama 3.1 8Bctx 131K
8.03Bparams
Llama 2 13Bctx 4K
13Bparams
Llama 4 Scout 17Bctx 131K
17Bparams
Llama 1 33Bctx 2K
33Bparams
Llama 1 65Bctx 2K
65Bparams
Llama 2 70Bctx 4K
70Bparams
Llama 3.3 70Bctx 131K
70Bparams
Llama 3 70Bctx 8K
70Bparams
Llama 3.1 70Bctx 131K
70.6Bparams
Llama 3.1 405Bctx 131K
405Bparams
Mistral 7B v0.1ctx 8K
7Bparams
Mistral 7B v0.3ctx 33K
7.24Bparams
Mistral Nemo 12Bctx 131K
12.2Bparams
Mixtral 8x7Bctx 33K
46.7Bparams
Mixtral 8x22Bctx 66K
141Bparams
Phi-2 2.7Bctx 2K
2.7Bparams
Phi-3.5 Mini 3.8Bctx 128K
3.8Bparams
Phi-3 Mini 3.8Bctx 128K
3.8Bparams
Phi-4 14Bctx 16K
14Bparams
Qwen 2.5 0.5Bctx 131K
0.5Bparams
Qwen 2.5 1.5Bctx 131K
1.5Bparams
Qwen 2.5 3Bctx 131K
3.1Bparams
Qwen 2 7Bctx 131K
7Bparams
Qwen 2.5 7Bctx 131K
7.6Bparams
Qwen 2.5 Coder 7Bctx 131K
7.6Bparams
Qwen 1.5 14Bctx 33K
14Bparams
Qwen 2.5 14Bctx 131K
14.7Bparams
Qwen 2.5 Coder 32Bctx 131K
32.5Bparams
Qwen 2.5 32Bctx 131K
32.5Bparams
Qwen 2 72Bctx 131K
72Bparams
Qwen 2.5 72Bctx 131K
72.7Bparams
Yi 34Bctx 200K
34Bparams

Quantization

F32
starstarstarstarstar
Maximum quality. Research only.
F16
starstarstarstarstar
No noticeable loss. Fine-tuning and inference.
Q8
starstarstarstarstar_border
Near identical to F16, half the VRAM.
Q6
starstarstarstarstar_border
Minimal loss. Good balance.
Q5
starstarstarstar_borderstar_border
Recommended if you have enough VRAM.
PopularQ4
starstarstarstar_borderstar_border
Most used. Best VRAM/quality balance.
Q3
starstarstar_borderstar_borderstar_border
Noticeable loss. Use only if necessary.
Q2
starstar_borderstar_borderstar_borderstar_border
Severe loss. Last resort.
SelectedQ4_K_M — Popular
Bits/weight4.5
QualityAcceptable

Context

Context length

VRAM by model

Frequently asked questions

add

How much VRAM does a language model need?

Roughly the parameter count multiplied by the bytes per parameter, plus the context. A 7-billion-parameter model at 4-bit needs about 4 GB, at 8-bit about 7 GB, and at full 16-bit precision about 14 GB. A 70B model at 4-bit needs around 40 GB, which is beyond any single consumer card.

add

What does quantization actually cost me?

It stores each weight in fewer bits, so the model shrinks roughly in proportion. Going from 16-bit to 8-bit is close to free in quality terms; 4-bit is the usual sweet spot and the loss is small for most uses; below 4-bit the degradation becomes noticeable.

add

Can I add a second graphics card to fit a bigger model?

Yes — the layers are split across cards and the totals add up, so two 16 GB cards can hold a model that needs 30 GB. Expect some loss to the traffic between them, and note that the slower card sets the pace.

add

What happens if the model does not fit?

Whatever does not fit spills into system memory, and the parts held there run at a fraction of the speed, because the processor's memory is far slower than the card's. A model that mostly fits is usable; one that mostly does not is not.

add

Does a longer context need more memory?

Yes, and people routinely forget it. The attention cache grows with the number of tokens in play, and at long contexts it can take more memory than the model weights themselves. Size for the context you intend to use, not the smallest one.

Keep going

Other tools and reference pages that pick up where this one leaves off.