Desktop GPU · NVIDIA
What AI models can a GeForce RTX 2060 12GB run?
12 GB is the point where local models stop being a demo. The popular 7B to 14B models fit with room for a real conversation, and a 32B fits if you accept a lower quantisation. At 336 GB/s there is enough bandwidth to keep generation responsive for anything that fits.
The short answer
Assuming an 8k context window and default settings, these are the models worth downloading first.
Best all-rounder
GPT OSS 20B
IQ4_XS · 11.13 GB · about 132.8 tokens/s
Best for code
DeepSeek Coder 6.7B Instruct
Q6_K · 10 GB · about 30.1 tokens/s
Largest that still runs well
GPT OSS 20B
20.91B parameters · IQ4_XS · about 132.8 tokens/s
Fastest useful answer
PowerMoE 3B
about 201 tokens/s · 4.53 GB
See how fast it feels
GeForce RTX 2060 12GB running GPT OSS 20B at IQ4_XS
YouWhy does my model use more memory when the conversation gets longer?
Because of the KV cache. Every token you send leaves behind a key and a value vector in each layer of the model, and those stay in memory for as long as the conversation lasts.
The weights are a fixed cost: load a 4-bit 8B model and that is about 4.8 GB, whether you write one word or ten thousand. The cache is the part that grows, and it grows in a straight line with the number of tokens in the window.
How steeply depends on the model's attention design. With grouped-query attention, several query heads share one key-value pair, which cuts the cache by that ratio. Without it, every head keeps its own, and a long context can cost more memory than the weights themselves.
If you are short on memory, the first thing to try is lowering the context window in your runtime, not the quantisation.
Reading your question
Simulated from our estimate at a 512-token question, not a recording. Assumes nothing else is competing for the GPU.
Every model, scored on this device
Each row uses the highest-quality quantisation that both fits and stays conversational. Speed is a single-stream estimate at 8k context.
| Model | Params | Verdict | Download | Memory used | Speed | Max context |
|---|---|---|---|---|---|---|
| GPT OSS 20B GPT-OSS | 20.91B (4.18B active) | Just fits | IQ4_XS | 11.13 GB | 132.8 t/s | 8k |
| Gemma 3 12B Instruct Gemma | 12.19B | Just fits | Q6_K | 10.96 GB | 27.2 t/s | 8k |
| Gemma 4 12B Instruct Gemma | 11.96B | Just fits | Q6_K | 10.34 GB | 29 t/s | 32k |
| vLLM Translategemma 12B Instruct Gemma | 13.19B | Runs great | Q5_K_M | 9.95 GB | 30.2 t/s | 32k |
| DeepSeek Coder V2 Lite Instruct DeepSeek | 15.71B (2.74B active) | Runs great | Q4_K_M | 9.8 GB | 155 t/s | 32k |
| Qwen2.5 14B Instruct Qwen | 14.77B | Just fits | Q4_K_M | 10.72 GB | 28.1 t/s | 8k |
| Qwen3 14B Qwen | 14.77B | Just fits | Q4_K_M | 10.47 GB | 28.8 t/s | 8k |
| Phi 4 Phi | 14.66B | Just fits | Q4_K_M | 10.72 GB | 28.1 t/s | 8k |
| Mistral Nemo Instruct 2407 Mistral | 12.25B | Just fits | Q5_K_M | 10.28 GB | 29.4 t/s | 8k |
| Qwen1.5 MoE A2.7B Qwen | 14.32B (2.69B active) | Just fits | Q4_K_M | 10.28 GB | 91.5 t/s | 8k |
| Qwen3.5 9B Qwen | 9.65B | Runs great | Q6_K | 9.22 GB | 32.9 t/s | 16k |
| Granite 4.1 8B Granite | 8.79B | Just fits | Q8_0 | 10.8 GB | 27.7 t/s | 8k |
| Internlm3 8B Instruct InternLM | 8.8B | Runs great | Q8_0 | 9.93 GB | 30.3 t/s | 16k |
| Gemma 2 9B Instruct Gemma | 9.24B | Runs great | Q6_K | 9.19 GB | 32.9 t/s | 8k |
| Fanar 1 9B Instruct Fanar | 8.78B | Just fits | Q8_0 | 10.82 GB | 27.6 t/s | 4k |
| LFM2.5 8B A1B Liquid | 8.47B (1.57B active) | Runs great | Q8_0 | 9.48 GB | 142.9 t/s | 32k |
| Qwen2.5 VL 7B Instruct Qwen | 8.29B | Runs great | Q8_0 | 9.46 GB | 31.9 t/s | 16k |
| Qwen3 8B Qwen | 8.19B | Runs great | Q8_0 | 10.08 GB | 29.9 t/s | 8k |
| Granite 3.0 8B Instruct Granite | 8.17B | Runs great | Q8_0 | 10.18 GB | 29.5 t/s | 4k |
| T Lite Instruct 2.1 T-Lite | 8.19B | Runs great | Q8_0 | 10.08 GB | 29.9 t/s | 8k |
| Llama 3.1 8B Instruct Llama | 8.03B | Runs great | Q8_0 | 9.8 GB | 30.8 t/s | 16k |
| Apertus 8B Instruct 2509 Apertus | 8.05B | Runs great | Q8_0 | 9.82 GB | 30.7 t/s | 16k |
| Llama 3 Taiwan 8B Instruct Llama | 8.03B | Runs great | Q8_0 | 9.8 GB | 30.8 t/s | 8k |
| Gemma 4 E4B Instruct Gemma | 8B | Runs great | Q8_0 | 8.72 GB | 34.6 t/s | 128k |
| Qwen2.5 7B Instruct Qwen | 7.62B | Runs great | Q8_0 | 8.8 GB | 34.5 t/s | 32k |
| Qwen1.5 7B Qwen | 7.72B | Just fits | Q6_K | 10.75 GB | 27.8 t/s | 8k |
| OLMo 3 7B Instruct OLMo | 7.3B | Runs great | Q8_0 | 10.07 GB | 29.9 t/s | 32k |
| Mistral 7B Instruct V0.3 Mistral | 7.25B | Runs great | Q8_0 | 9.02 GB | 33.7 t/s | 16k |
| Mistral 7B Instruct V0.2 Mistral | 7.24B | Runs great | Q8_0 | 9.01 GB | 33.7 t/s | 16k |
| Gemma 4 26B A4B Instruct Gemma | 25.81B | Just fits | IQ3_XXS ! | 10.2 GB | 29.2 t/s | 32k |
| OLMoE 1B 7B 0125 Instruct OLMo | 6.92B (1.28B active) | Runs great | Q8_0 | 8.57 GB | 121.6 t/s | 4k |
| Falcon 7B Falcon | 7.22B | Just fits | Q6_K | 10.83 GB | 27.7 t/s | 8k |
| Mistral Small 24B Instruct 2501 Mistral | 23.57B | Just fits | IQ3_XXS ! | 10.56 GB | 28.6 t/s | 8k |
| DeepSeek Coder 7B Instruct V1.5 DeepSeek | 6.91B | Runs great | Q6_K | 9.88 GB | 30.5 t/s | 4k |
| CodeLlama 7B Llama | 6.74B | Runs great | Q6_K | 10 GB | 30.1 t/s | 8k |
| DeepSeek Coder 6.7B Instruct DeepSeek | 6.74B | Runs great | Q6_K | 10 GB | 30.1 t/s | 8k |
| Codestral 22B V0.1 Mistral | 22.25B | Just fits | IQ3_XXS ! | 10.65 GB | 28.5 t/s | 8k |
| Gemma 4 E2B Instruct Gemma | 5.12B | Runs great | Q8_0 | 5.78 GB | 54.2 t/s | 128k |
| Qwen3.5 4B Qwen | 4.66B | Runs great | Q8_0 | 6.37 GB | 49.1 t/s | 32k |
| Agents A1 4B Other | 4.54B | Runs great | Q8_0 | 6.25 GB | 50.2 t/s | 32k |
| Gemma 3 4B Instruct Gemma | 4.3B | Runs great | Q8_0 | 5.31 GB | 60.5 t/s | 128k |
| Phi 3 Vision 128k Instruct Phi | 4.15B | Runs great | Q8_0 | 7.89 GB | 38.8 t/s | 16k |
| Qwen3 4B Qwen | 4.02B | Runs great | Q8_0 | 5.86 GB | 54 t/s | 32k |
| Phi 4 Mini Instruct Phi | 3.84B | Runs great | Q8_0 | 5.59 GB | 57.4 t/s | 32k |
| Phi 3 Mini 4k Instruct Phi | 3.82B | Runs great | Q8_0 | 5.32 GB | 60.8 t/s | 4k |
| PowerLM 3B PowerLM | 3.51B | Runs great | Q8_0 | 7.03 GB | 43.8 t/s | 4k |
| Granite 4.1 3B Granite | 3.4B | Runs great | Q8_0 | 4.75 GB | 69.1 t/s | 64k |
| PowerMoE 3B PowerLM | 3.37B (0.88B active) | Runs great | Q8_0 | 4.53 GB | 201 t/s | 4k |
| Llama 3.2 3B Instruct Llama | 3.21B | Runs great | Q8_0 | 4.84 GB | 68 t/s | 32k |
| Qwen3.5 35B A3B Qwen | 35.95B (2.9B active) | Runs great | IQ2_XXS ! | 9.97 GB | 208.7 t/s | 16k |
| Qwen2.5 3B Instruct Qwen | 3.09B | Runs great | Q8_0 | 4.07 GB | 82.5 t/s | 32k |
| SmolLM3 3B Base SmolLM | 3.08B | Runs great | Q8_0 | 4.34 GB | 76.3 t/s | 64k |
| Starcoder2 3B StarCoder | 3.03B | Runs great | Q8_0 | 3.9 GB | 88.4 t/s | 16k |
| Qwen3 32B Qwen | 32.76B | Just fits | IQ2_XXS ! | 10.77 GB | 28 t/s | 8k |
| Qwen2.5 32B Instruct Qwen | 32.76B | Just fits | IQ2_XXS ! | 10.77 GB | 28 t/s | 8k |
| OTel 2.0 LLM 31B Instruct Other | 32.11B | Runs great | IQ2_XXS ! | 9.57 GB | 31.9 t/s | 32k |
| Gemma 4 31B Instruct Gemma | 31.27B | Runs great | IQ2_XXS ! | 9.37 GB | 32.7 t/s | 32k |
| GLM 4.7 Flash GLM | 31.22B (3.66B active) | Runs great | IQ2_XXS ! | 8.63 GB | 213.4 t/s | 32k |
| Qwen3 30B A3B Qwen | 30.53B (3.34B active) | Runs great | IQ2_XXS ! | 8.8 GB | 177.6 t/s | 16k |
| Granite 4.1 30B Granite | 28.87B | Runs great | IQ2_XXS ! | 9.77 GB | 30.9 t/s | 8k |
| Phi 2 Phi | 2.78B | Runs great | Q8_0 | 6.01 GB | 52.5 t/s | 2k |
| Qwen3.5 27B Qwen | 27.78B | Runs great | IQ2_XXS ! | 9.58 GB | 31.8 t/s | 8k |
| Gemma 3 27B Instruct Gemma | 27.43B | Runs great | IQ2_XXS ! | 8.59 GB | 35.9 t/s | 16k |
| LFM2.5 2.6B Liquid | 2.7B | Runs great | Q8_0 | 3.87 GB | 87.7 t/s | 64k |
| Gemma 2 2B Instruct Gemma | 2.61B | Runs great | Q8_0 | 3.73 GB | 92.2 t/s | 8k |
| Qwen3.5 2B Qwen | 2.27B | Runs great | Q8_0 | 3.35 GB | 105.1 t/s | 128k |
| OneRec 1.7B Other | 2.13B | Runs great | Q8_0 | 3.71 GB | 92.4 t/s | 32k |
| Qwen3 1.7B Qwen | 2.03B | Runs great | Q8_0 | 3.61 GB | 95.5 t/s | 32k |
| OLMo 2 1124 13B Instruct OLMo | 13.72B | Just fits | IQ2_XXS ! | 10.45 GB | 28.9 t/s | 4k |
| DeepSeek R1 Distill Qwen 1.5B Qwen | 1.78B | Runs great | Q8_0 | 2.57 GB | 147.3 t/s | 128k |
| Qwen3 1.7B Base Qwen | 1.72B | Runs great | Q8_0 | 3.3 GB | 106.9 t/s | 32k |
| SmolLM2 1.7B SmolLM | 1.71B | Runs great | Q8_0 | 3.92 GB | 86.3 t/s | 8k |
| Qwen2.5 1.5B Instruct Qwen | 1.54B | Runs great | Q8_0 | 2.44 GB | 158.1 t/s | 32k |
| Pythia 1.4B Pythia | 1.52B | Runs great | Q8_0 | 3.73 GB | 91.7 t/s | 2k |
| OLMo 2 0425 1B OLMo | 1.48B | Runs great | Q8_0 | 3.19 GB | 111.8 t/s | 4k |
| Llama 3.2 1B Instruct Llama | 1.24B | Runs great | Q8_0 | 2.2 GB | 186.5 t/s | 128k |
| LFM2.5 1.2B Instruct Liquid | 1.17B | Runs great | Q8_0 | 2.13 GB | 195.7 t/s | 64k |
| TinyLlama 1.1B Chat V1.0 Llama | 1.1B | Runs great | Q8_0 | 1.99 GB | 218.6 t/s | 2k |
| MiniCPM5 1B MiniCPM | 1.08B | Runs great | Q8_0 | 1.95 GB | 219.3 t/s | 128k |
| Gemma 3 1B Instruct Gemma | 1B | Runs great | Q8_0 | 1.71 GB | 265.2 t/s | 32k |
| Qwen3.5 0.8B Qwen | 0.87B | Runs great | Q8_0 | 1.9 GB | 222.9 t/s | 128k |
| Sarashina2.2 0.5B Instruct V0.1 Sarashina | 0.79B | Runs great | Q8_0 | 1.93 GB | 220.3 t/s | 8k |
| Qwen3 0.6B Qwen | 0.75B | Runs great | Q8_0 | 2.28 GB | 170.4 t/s | 32k |
| Qwen1.5 0.5B Chat Qwen | 0.62B | Runs great | Q8_0 | 2.03 GB | 202.1 t/s | 32k |
| Qwen3 0.6B Base Qwen | 0.6B | Runs great | Q8_0 | 2.13 GB | 187.6 t/s | 32k |
| Pythia 410m Pythia | 0.51B | Runs great | Q8_0 | 1.92 GB | 219.6 t/s | 2k |
| H2o Danube3 500m Chat Danube | 0.51B | Runs great | Q8_0 | 1.57 GB | 313.2 t/s | 8k |
| Qwen2.5 0.5B Instruct Qwen | 0.49B | Runs great | Q8_0 | 1.23 GB | 476.2 t/s | 32k |
| SmolLM2 360M SmolLM | 0.36B | Runs great | Q8_0 | 1.33 GB | 412 t/s | 8k |
| LFM2.5 350M Liquid | 0.35B | Runs great | Q8_0 | 1.26 GB | 462 t/s | 64k |
Showing the 90 best results of 133. The remaining 43 need more memory than this device has, at any quantisation.
What it will not run
2 of the 133 architectures we track are out of reach here, even at two-bit precision. Another 35 run by splitting layers between the card and system RAM, which works but drops generation to single digits.
If you want more headroom
The next steps up in memory, in order. More memory changes which models load at all; more bandwidth changes how fast they answer, so the two are worth weighing separately.
| Device | Memory | Bandwidth | Models that fit | |
|---|---|---|---|---|
| GeForce RTX 5080 | 16 GB | 960 GB/s | 99 | Check price |
| GeForce RTX 5070 Ti | 16 GB | 896 GB/s | 99 | Check price |
| GeForce RTX 4080 SUPER | 16 GB | 736 GB/s | 99 | Check price |
| GeForce RTX 4080 | 16 GB | 717 GB/s | 99 | Check price |
Price links go to an Amazon search for the model name and are affiliate links: if you buy through one, we earn a commission at no cost to you. We do not take payment for placement, and the ordering above is by memory capacity alone.
Other devices with 12 GB
Same capacity, different speed. Once a model fits, bandwidth is what separates these.