Desktop GPU · NVIDIA
What AI models can a GeForce GTX 1650 run?
With 4 GB to work with, the question is not which model is best but which ones fit at all. The answer is: small ones, and they are better than they used to be. Bandwidth is the weak point at 128 GB/s. Models fit, then generate slowly, because every token means reading the whole active model out of memory.
The short answer
Assuming an 8k context window and default settings, these are the models worth downloading first.
See how fast it feels
GeForce GTX 1650 running Gemma 3 4B Instruct at IQ4_XS
YouWhy does my model use more memory when the conversation gets longer?
Because of the KV cache. Every token you send leaves behind a key and a value vector in each layer of the model, and those stay in memory for as long as the conversation lasts.
The weights are a fixed cost: load a 4-bit 8B model and that is about 4.8 GB, whether you write one word or ten thousand. The cache is the part that grows, and it grows in a straight line with the number of tokens in the window.
How steeply depends on the model's attention design. With grouped-query attention, several query heads share one key-value pair, which cuts the cache by that ratio. Without it, every head keeps its own, and a long context can cost more memory than the weights themselves.
If you are short on memory, the first thing to try is lowering the context window in your runtime, not the quantisation.
Reading your question
Simulated from our estimate at a 512-token question, not a recording. Assumes nothing else is competing for the GPU.
Every model, scored on this device
Each row uses the highest-quality quantisation that both fits and stays conversational. Speed is a single-stream estimate at 8k context.
| Model | Params | Verdict | Download | Memory used | Speed | Max context |
|---|---|---|---|---|---|---|
| Gemma 3 4B Instruct Gemma | 4.3B | Just fits | IQ4_XS | 3.18 GB | 43.3 t/s | 8k |
| Gemma 4 E2B Instruct Gemma | 5.12B | Just fits | Q3_K_M | 3.04 GB | 44.7 t/s | 16k |
| Qwen2.5 3B Instruct Qwen | 3.09B | Just fits | Q5_K_M | 3.05 GB | 45.1 t/s | 8k |
| Starcoder2 3B StarCoder | 3.03B | Just fits | Q5_K_M | 2.91 GB | 49.4 t/s | 16k |
| PowerMoE 3B PowerLM | 3.37B (0.88B active) | Just fits | Q4_K_M | 3.09 GB | 105.5 t/s | 4k |
| Gemma 2 2B Instruct Gemma | 2.61B | Just fits | Q6_K | 3.14 GB | 43.7 t/s | 8k |
| Granite 4.1 3B Granite | 3.4B | Just fits | IQ4_XS | 3.06 GB | 45.5 t/s | 8k |
| SmolLM3 3B Base SmolLM | 3.08B | Just fits | Q4_K_M | 3.02 GB | 45.7 t/s | 8k |
| LFM2.5 2.6B Liquid | 2.7B | Just fits | Q5_K_M | 2.98 GB | 46.5 t/s | 8k |
| Qwen3.5 2B Qwen | 2.27B | Just fits | Q6_K | 2.84 GB | 49.8 t/s | 8k |
| Qwen3 1.7B Qwen | 2.03B | Just fits | Q6_K | 3.15 GB | 43.3 t/s | 8k |
| Llama 3.2 3B Instruct Llama | 3.21B | Just fits | Q3_K_M | 3.12 GB | 44.9 t/s | 8k |
| OneRec 1.7B Other | 2.13B | Just fits | Q5_K_M | 3.01 GB | 45.9 t/s | 8k |
| DeepSeek R1 Distill Qwen 1.5B Qwen | 1.78B | Just fits | Q8_0 | 2.57 GB | 56.1 t/s | 32k |
| Phi 4 Mini Instruct Phi | 3.84B | Just fits | IQ3_XXS ! | 3.16 GB | 44.3 t/s | 8k |
| Phi 3 Mini 4k Instruct Phi | 3.82B | Just fits | IQ3_XXS ! | 2.9 GB | 49.7 t/s | 4k |
| Qwen3 1.7B Base Qwen | 1.72B | Just fits | Q6_K | 2.91 GB | 48 t/s | 8k |
| Qwen2.5 1.5B Instruct Qwen | 1.54B | Just fits | Q8_0 | 2.44 GB | 60.2 t/s | 16k |
| OLMo 2 0425 1B OLMo | 1.48B | Just fits | Q8_0 | 3.19 GB | 42.6 t/s | 4k |
| SmolLM2 1.7B SmolLM | 1.71B | Just fits | Q4_K_M | 3.19 GB | 42.6 t/s | 8k |
| LFM2.5 8B A1B Liquid | 8.47B (1.57B active) | Just fits | IQ2_XXS ! | 3.13 GB | 139.7 t/s | 8k |
| Gemma 4 E4B Instruct Gemma | 8B | Just fits | IQ2_XXS ! | 2.72 GB | 53.6 t/s | 32k |
| Qwen2.5 7B Instruct Qwen | 7.62B | Just fits | IQ2_XXS ! | 3.08 GB | 46.3 t/s | 8k |
| Pythia 1.4B Pythia | 1.52B | Just fits | Q4_K_M | 3.08 GB | 44.6 t/s | 2k |
| Llama 3.2 1B Instruct Llama | 1.24B | Just fits | Q8_0 | 2.2 GB | 71.1 t/s | 16k |
| LFM2.5 1.2B Instruct Liquid | 1.17B | Runs great | Q8_0 | 2.13 GB | 74.6 t/s | 16k |
| TinyLlama 1.1B Chat V1.0 Llama | 1.1B | Runs great | Q8_0 | 1.99 GB | 83.3 t/s | 2k |
| MiniCPM5 1B MiniCPM | 1.08B | Runs great | Q8_0 | 1.95 GB | 83.6 t/s | 32k |
| Gemma 3 1B Instruct Gemma | 1B | Runs great | Q8_0 | 1.71 GB | 101 t/s | 32k |
| Qwen3.5 4B Qwen | 4.66B | Just fits | IQ2_XXS ! | 2.88 GB | 49.6 t/s | 8k |
| Agents A1 4B Other | 4.54B | Just fits | IQ2_XXS ! | 2.85 GB | 50.2 t/s | 8k |
| Qwen3 4B Qwen | 4.02B | Just fits | IQ2_XXS ! | 2.85 GB | 50.2 t/s | 8k |
| Qwen3.5 0.8B Qwen | 0.87B | Runs great | Q8_0 | 1.9 GB | 84.9 t/s | 16k |
| Sarashina2.2 0.5B Instruct V0.1 Sarashina | 0.79B | Runs great | Q8_0 | 1.93 GB | 83.9 t/s | 8k |
| Qwen3 0.6B Qwen | 0.75B | Just fits | Q8_0 | 2.28 GB | 64.9 t/s | 8k |
| Qwen1.5 0.5B Chat Qwen | 0.62B | Runs great | Q8_0 | 2.03 GB | 77 t/s | 16k |
| Qwen3 0.6B Base Qwen | 0.6B | Runs great | Q8_0 | 2.13 GB | 71.5 t/s | 16k |
| Pythia 410m Pythia | 0.51B | Runs great | Q8_0 | 1.92 GB | 83.7 t/s | 2k |
| H2o Danube3 500m Chat Danube | 0.51B | Runs great | Q8_0 | 1.57 GB | 119.3 t/s | 8k |
| Qwen2.5 0.5B Instruct Qwen | 0.49B | Runs great | Q8_0 | 1.23 GB | 181.4 t/s | 32k |
| SmolLM2 360M SmolLM | 0.36B | Runs great | Q8_0 | 1.33 GB | 157 t/s | 8k |
| LFM2.5 350M Liquid | 0.35B | Runs great | Q8_0 | 1.26 GB | 176 t/s | 32k |
| Pythia 160m Pythia | 0.21B | Runs great | Q8_0 | 1.14 GB | 214.6 t/s | 2k |
| Japanese GPT NeoX Small GPT-NeoX | 0.2B | Runs great | Q8_0 | 1.13 GB | 219.1 t/s | 2k |
| Llama 160m Llama | 0.16B | Runs great | Q8_0 | 1.09 GB | 238.8 t/s | 2k |
| LLM Jp 3 150m LLM-jp | 0.15B | Runs great | Q8_0 | 0.97 GB | 312.4 t/s | 4k |
| SmolLM2 135M SmolLM | 0.13B | Runs great | Q8_0 | 0.94 GB | 344.8 t/s | 8k |
| Pythia 70m Deduped Pythia | 0.1B | Runs great | Q8_0 | 0.82 GB | 544.7 t/s | 2k |
| OLMoE 1B 7B 0125 Instruct OLMo | 6.92B (1.28B active) | Partial offload 8/16 layers | — | 5.62 GB | 28.3 t/s | — |
| DeepSeek Coder V2 Lite Instruct DeepSeek | 15.71B (2.74B active) | Partial offload 7/27 layers | — | 9.8 GB | 21.7 t/s | — |
| GPT OSS 20B GPT-OSS | 20.91B (4.18B active) | Partial offload 4/24 layers | — | 12.54 GB | 15.2 t/s | — |
| Qwen3.5 35B A3B Qwen | 35.95B (2.9B active) | Partial offload 4/40 layers | — | 21.57 GB | 15.1 t/s | — |
| GLM 4.7 Flash GLM | 31.22B (3.66B active) | Partial offload 6/47 layers | — | 18.69 GB | 14 t/s | — |
| Phi 2 Phi | 2.78B | Partial offload 19/32 layers | — | 4.82 GB | 13.3 t/s | — |
| Qwen3 30B A3B Qwen | 30.53B (3.34B active) | Partial offload 6/48 layers | — | 18.64 GB | 13.2 t/s | — |
| Qwen3 Next 80B A3B Instruct Qwen | 81.32B (3.19B active) | Partial offload 2/48 layers | — | 47.2 GB | 12.8 t/s | — |
| Qwen3 Coder Next Qwen | 79.67B (3.19B active) | Partial offload 2/48 layers | — | 46.27 GB | 12.8 t/s | — |
| Qwen1.5 MoE A2.7B Qwen | 14.32B (2.69B active) | Partial offload 6/24 layers | — | 10.28 GB | 12.7 t/s | — |
| PowerLM 3B PowerLM | 3.51B | Partial offload 20/40 layers | — | 5.53 GB | 10.2 t/s | — |
| GPT OSS 120B GPT-OSS | 116.83B (5.7B active) | Partial offload 1/36 layers | — | 66.48 GB | 10 t/s | — |
| Qwen2.5 VL 7B Instruct Qwen | 8.29B | Partial offload 13/28 layers | — | 5.92 GB | 9.2 t/s | — |
| Mistral 7B Instruct V0.3 Mistral | 7.25B | Partial offload 14/32 layers | — | 5.93 GB | 9 t/s | — |
| Mistral 7B Instruct V0.2 Mistral | 7.24B | Partial offload 14/32 layers | — | 5.92 GB | 9 t/s | — |
| Internlm3 8B Instruct InternLM | 8.8B | Partial offload 21/48 layers | — | 6.17 GB | 8.6 t/s | — |
| Phi 3 Vision 128k Instruct Phi | 4.15B | Partial offload 14/32 layers | — | 6.12 GB | 8.5 t/s | — |
| Apertus 8B Instruct 2509 Apertus | 8.05B | Partial offload 13/32 layers | — | 6.38 GB | 8 t/s | — |
| Llama 3.1 8B Instruct Llama | 8.03B | Partial offload 13/32 layers | — | 6.37 GB | 8 t/s | — |
| Llama 3 Taiwan 8B Instruct Llama | 8.03B | Partial offload 13/32 layers | — | 6.37 GB | 8 t/s | — |
| Qwen3 8B Qwen | 8.19B | Partial offload 14/36 layers | — | 6.58 GB | 7.6 t/s | — |
| T Lite Instruct 2.1 T-Lite | 8.19B | Partial offload 14/36 layers | — | 6.58 GB | 7.6 t/s | — |
| Granite 3.0 8B Instruct Granite | 8.17B | Partial offload 16/40 layers | — | 6.69 GB | 7.5 t/s | — |
| Laguna S 2.1 Laguna | 117.56B (7.71B active) | Partial offload 1/48 layers | — | 66.98 GB | 7.2 t/s | — |
| Phi 3.5 MoE Instruct Phi | 41.87B (6.64B active) | Partial offload 3/32 layers | — | 25.39 GB | 7.1 t/s | — |
| OLMo 3 7B Instruct OLMo | 7.3B | Partial offload 12/32 layers | — | 6.96 GB | 7 t/s | — |
| Granite 4.1 8B Granite | 8.79B | Partial offload 15/40 layers | — | 7.04 GB | 6.9 t/s | — |
| Fanar 1 9B Instruct Fanar | 8.78B | Partial offload 15/42 layers | — | 7.07 GB | 6.7 t/s | — |
| Qwen3.5 9B Qwen | 9.65B | Partial offload 11/32 layers | — | 7.28 GB | 6.5 t/s | — |
| Gemma 2 9B Instruct Gemma | 9.24B | Partial offload 15/42 layers | — | 7.33 GB | 6.5 t/s | — |
| Qwen3.5 122B A10B Qwen | 125.09B (8.17B active) | Partial offload 1/48 layers | — | 71.88 GB | 6 t/s | — |
| Gemma 4 12B Instruct Gemma | 11.96B | Partial offload 15/48 layers | — | 7.94 GB | 5.7 t/s | — |
| Gemma 3 12B Instruct Gemma | 12.19B | Partial offload 14/48 layers | — | 8.5 GB | 5.2 t/s | — |
| DeepSeek Coder 7B Instruct V1.5 DeepSeek | 6.91B | Partial offload 9/30 layers | — | 8.49 GB | 5.2 t/s | — |
| vLLM Translategemma 12B Instruct Gemma | 13.19B | Partial offload 14/48 layers | — | 8.63 GB | 5.1 t/s | — |
| DeepSeek Coder 6.7B Instruct DeepSeek | 6.74B | Partial offload 9/32 layers | — | 8.64 GB | 5.1 t/s | — |
| CodeLlama 7B Llama | 6.74B | Partial offload 9/32 layers | — | 8.64 GB | 5.1 t/s | — |
| Mistral Nemo Instruct 2407 Mistral | 12.25B | Partial offload 11/40 layers | — | 9.05 GB | 4.8 t/s | — |
| Qwen1.5 7B Qwen | 7.72B | Partial offload 9/32 layers | — | 9.19 GB | 4.7 t/s | — |
| Falcon 7B Falcon | 7.22B | Partial offload 8/32 layers | — | 9.38 GB | 4.5 t/s | — |
| Mixtral 8x7B Instruct V0.1 Mistral | 46.7B (12.88B active) | Partial offload 2/32 layers | — | 28.11 GB | 4 t/s | — |
| MiniMax M2.5 MiniMax | 228.7B (10.98B active) | Partial offload 1/62 layers | — | 131.32 GB | 3.9 t/s | — |
Showing the 90 best results of 133. The remaining 43 need more memory than this device has, at any quantisation.
What it will not run
18 of the 133 architectures we track are out of reach here, even at two-bit precision. Another 67 run by splitting layers between the card and system RAM, which works but drops generation to single digits.
If you want more headroom
The next steps up in memory, in order. More memory changes which models load at all; more bandwidth changes how fast they answer, so the two are worth weighing separately.
| Device | Memory | Bandwidth | Models that fit | |
|---|---|---|---|---|
| GeForce GTX 1660 SUPER | 6 GB | 336 GB/s | 72 | Check price |
| GeForce RTX 3050 6GB | 6 GB | 168 GB/s | 72 | Check price |
| GeForce RTX 3070 Ti | 8 GB | 608 GB/s | 83 | Check price |
| Arc A750 | 8 GB | 512 GB/s | 83 | Check price |
Price links go to an Amazon search for the model name and are affiliate links: if you buy through one, we earn a commission at no cost to you. We do not take payment for placement, and the ordering above is by memory capacity alone.