Desktop GPU · NVIDIA

What AI models can a GeForce RTX 3050 6GB run?

With 6 GB to work with, the question is not which model is best but which ones fit at all. The answer is: small ones, and they are better than they used to be. Bandwidth is the weak point at 168 GB/s. Models fit, then generate slowly, because every token means reading the whole active model out of memory.

Memory6 GB
Bandwidth168 GB/s
Usable for a model5.2 GB
Runtime backendCUDA
ArchitectureAmpere
Power70 W

The short answer

Assuming an 8k context window and default settings, these are the models worth downloading first.

See how fast it feels

GeForce RTX 3050 6GB running Gemma 4 E4B Instruct at IQ4_XS

Wait for the first word770 ms
Then writes at34.4 tok/s
Whole answer5.8 s

YouWhy does my model use more memory when the conversation gets longer?

Model

Because of the KV cache. Every token you send leaves behind a key and a value vector in each layer of the model, and those stay in memory for as long as the conversation lasts.

The weights are a fixed cost: load a 4-bit 8B model and that is about 4.8 GB, whether you write one word or ten thousand. The cache is the part that grows, and it grows in a straight line with the number of tokens in the window.

How steeply depends on the model's attention design. With grouped-query attention, several query heads share one key-value pair, which cuts the cache by that ratio. Without it, every head keeps its own, and a long context can cost more memory than the weights themselves.

If you are short on memory, the first thing to try is lowering the context window in your runtime, not the quantisation.

Simulated from our estimate at a 512-token question, not a recording. Assumes nothing else is competing for the GPU.

Every model, scored on this device

Each row uses the highest-quality quantisation that both fits and stays conversational. Speed is a single-stream estimate at 8k context.

Model Params Verdict Download Memory used Speed Max context
Gemma 4 E4B Instruct
Gemma
8B Just fits IQ4_XS 4.76 GB 34.4 t/s 16k
Qwen2.5 7B Instruct
Qwen
7.62B Just fits IQ4_XS 5.03 GB 32.7 t/s 8k
Gemma 4 E2B Instruct
Gemma
5.12B Just fits Q6_K 4.62 GB 35.1 t/s 32k
OLMoE 1B 7B 0125 Instruct
OLMo
6.92B (1.28B active) Just fits IQ4_XS 5.15 GB 84.3 t/s 4k
LFM2.5 8B A1B
Liquid
8.47B (1.57B active) Just fits Q3_K_M 4.96 GB 126.4 t/s 8k
Qwen2.5 VL 7B Instruct
Qwen
8.29B Just fits Q3_K_M 5.03 GB 32.7 t/s 8k
Gemma 3 4B Instruct
Gemma
4.3B Just fits Q6_K 4.34 GB 38.5 t/s 16k
Qwen3.5 4B
Qwen
4.66B Just fits Q5_K_M 4.84 GB 33.7 t/s 8k
Agents A1 4B
Other
4.54B Just fits Q5_K_M 4.77 GB 34.4 t/s 8k
Mistral 7B Instruct V0.3
Mistral
7.25B Just fits Q3_K_M 5.15 GB 32 t/s 8k
Mistral 7B Instruct V0.2
Mistral
7.24B Just fits Q3_K_M 5.15 GB 32.1 t/s 8k
Qwen3 4B
Qwen
4.02B Just fits Q6_K 4.95 GB 32.8 t/s 8k
Phi 4 Mini Instruct
Phi
3.84B Just fits Q6_K 4.72 GB 35 t/s 8k
Phi 3 Mini 4k Instruct
Phi
3.82B Just fits Q6_K 4.45 GB 37.6 t/s 4k
Granite 4.1 3B
Granite
3.4B Just fits Q8_0 4.75 GB 34.5 t/s 8k
PowerMoE 3B
PowerLM
3.37B (0.88B active) Just fits Q8_0 4.53 GB 100.5 t/s 4k
Internlm3 8B Instruct
InternLM
8.8B Just fits IQ3_XXS ! 4.36 GB 39.2 t/s 16k
Llama 3.2 3B Instruct
Llama
3.21B Just fits Q8_0 4.84 GB 34 t/s 8k
Qwen3 8B
Qwen
8.19B Just fits IQ3_XXS ! 4.89 GB 34.1 t/s 8k
Granite 3.0 8B Instruct
Granite
8.17B Just fits IQ3_XXS ! 5.01 GB 33.1 t/s 4k
T Lite Instruct 2.1
T-Lite
8.19B Just fits IQ3_XXS ! 4.89 GB 34.1 t/s 8k
Qwen2.5 3B Instruct
Qwen
3.09B Runs great Q8_0 4.07 GB 41.3 t/s 16k
SmolLM3 3B Base
SmolLM
3.08B Just fits Q8_0 4.34 GB 38.2 t/s 16k
Llama 3.1 8B Instruct
Llama
8.03B Just fits IQ3_XXS ! 4.71 GB 35.7 t/s 8k
Apertus 8B Instruct 2509
Apertus
8.05B Just fits IQ3_XXS ! 4.72 GB 35.6 t/s 8k
Llama 3 Taiwan 8B Instruct
Llama
8.03B Just fits IQ3_XXS ! 4.71 GB 35.7 t/s 8k
Starcoder2 3B
StarCoder
3.03B Runs great Q8_0 3.9 GB 44.2 t/s 16k
LFM2.5 2.6B
Liquid
2.7B Runs great Q8_0 3.87 GB 43.9 t/s 16k
Gemma 2 2B Instruct
Gemma
2.61B Runs great Q8_0 3.73 GB 46.1 t/s 8k
Phi 2
Phi
2.78B Just fits Q5_K_M 5.1 GB 31.7 t/s 2k
Qwen3.5 2B
Qwen
2.27B Runs great Q8_0 3.35 GB 52.6 t/s 32k
OneRec 1.7B
Other
2.13B Runs great Q8_0 3.71 GB 46.2 t/s 16k
PowerLM 3B
PowerLM
3.51B Just fits Q3_K_M 5.15 GB 31.2 t/s 4k
Qwen3 1.7B
Qwen
2.03B Runs great Q8_0 3.61 GB 47.8 t/s 16k
DeepSeek Coder V2 Lite Instruct
DeepSeek
15.71B (2.74B active) Just fits IQ2_XXS ! 4.73 GB 154 t/s 16k
vLLM Translategemma 12B Instruct
Gemma
13.19B Just fits IQ2_XXS ! 4.37 GB 38.9 t/s 32k
DeepSeek R1 Distill Qwen 1.5B
Qwen
1.78B Runs great Q8_0 2.57 GB 73.6 t/s 128k
Gemma 3 12B Instruct
Gemma
12.19B Just fits IQ2_XXS ! 4.57 GB 36.9 t/s 8k
Mistral Nemo Instruct 2407
Mistral
12.25B Just fits IQ2_XXS ! 5.1 GB 32.9 t/s 8k
Gemma 4 12B Instruct
Gemma
11.96B Runs great IQ2_XXS ! 4.08 GB 42.5 t/s 32k
Qwen3 1.7B Base
Qwen
1.72B Runs great Q8_0 3.3 GB 53.5 t/s 16k
SmolLM2 1.7B
SmolLM
1.71B Runs great Q8_0 3.92 GB 43.2 t/s 8k
Qwen2.5 1.5B Instruct
Qwen
1.54B Runs great Q8_0 2.44 GB 79.1 t/s 32k
Qwen3.5 9B
Qwen
9.65B Runs great IQ2_XXS ! 4.17 GB 41.6 t/s 8k
Pythia 1.4B
Pythia
1.52B Runs great Q8_0 3.73 GB 45.9 t/s 2k
Gemma 2 9B Instruct
Gemma
9.24B Just fits IQ2_XXS ! 4.35 GB 39 t/s 8k
OLMo 2 0425 1B
OLMo
1.48B Runs great Q8_0 3.19 GB 55.9 t/s 4k
Granite 4.1 8B
Granite
8.79B Just fits IQ2_XXS ! 4.21 GB 41 t/s 8k
Fanar 1 9B Instruct
Fanar
8.78B Just fits IQ2_XXS ! 4.24 GB 40.3 t/s 4k
OLMo 3 7B Instruct
OLMo
7.3B Just fits IQ2_XXS ! 4.6 GB 36.7 t/s 16k
Llama 3.2 1B Instruct
Llama
1.24B Runs great Q8_0 2.2 GB 93.3 t/s 64k
LFM2.5 1.2B Instruct
Liquid
1.17B Runs great Q8_0 2.13 GB 97.9 t/s 64k
TinyLlama 1.1B Chat V1.0
Llama
1.1B Runs great Q8_0 1.99 GB 109.3 t/s 2k
MiniCPM5 1B
MiniCPM
1.08B Runs great Q8_0 1.95 GB 109.7 t/s 64k
Gemma 3 1B Instruct
Gemma
1B Runs great Q8_0 1.71 GB 132.6 t/s 32k
Phi 3 Vision 128k Instruct
Phi
4.15B Just fits IQ2_XXS ! 4.78 GB 34.5 t/s 8k
Qwen3.5 0.8B
Qwen
0.87B Runs great Q8_0 1.9 GB 111.5 t/s 64k
Sarashina2.2 0.5B Instruct V0.1
Sarashina
0.79B Runs great Q8_0 1.93 GB 110.2 t/s 8k
Qwen3 0.6B
Qwen
0.75B Runs great Q8_0 2.28 GB 85.2 t/s 32k
Qwen1.5 0.5B Chat
Qwen
0.62B Runs great Q8_0 2.03 GB 101 t/s 32k
Qwen3 0.6B Base
Qwen
0.6B Runs great Q8_0 2.13 GB 93.8 t/s 32k
Pythia 410m
Pythia
0.51B Runs great Q8_0 1.92 GB 109.8 t/s 2k
H2o Danube3 500m Chat
Danube
0.51B Runs great Q8_0 1.57 GB 156.6 t/s 8k
Qwen2.5 0.5B Instruct
Qwen
0.49B Runs great Q8_0 1.23 GB 238.1 t/s 32k
SmolLM2 360M
SmolLM
0.36B Runs great Q8_0 1.33 GB 206 t/s 8k
LFM2.5 350M
Liquid
0.35B Runs great Q8_0 1.26 GB 231 t/s 64k
Pythia 160m
Pythia
0.21B Runs great Q8_0 1.14 GB 281.7 t/s 2k
Japanese GPT NeoX Small
GPT-NeoX
0.2B Runs great Q8_0 1.13 GB 287.5 t/s 2k
Llama 160m
Llama
0.16B Runs great Q8_0 1.09 GB 313.4 t/s 2k
LLM Jp 3 150m
LLM-jp
0.15B Runs great Q8_0 0.97 GB 410.1 t/s 4k
SmolLM2 135M
SmolLM
0.13B Runs great Q8_0 0.94 GB 452.5 t/s 8k
Pythia 70m Deduped
Pythia
0.1B Runs great Q8_0 0.82 GB 714.9 t/s 2k
GPT OSS 20B
GPT-OSS
20.91B (4.18B active) Partial offload 9/24 layers 12.54 GB 18.9 t/s
Qwen3.5 35B A3B
Qwen
35.95B (2.9B active) Partial offload 8/40 layers 21.57 GB 16.6 t/s
Qwen1.5 MoE A2.7B
Qwen
14.32B (2.69B active) Partial offload 11/24 layers 10.28 GB 16.2 t/s
GLM 4.7 Flash
GLM
31.22B (3.66B active) Partial offload 11/47 layers 18.69 GB 15.6 t/s
Qwen3 30B A3B
Qwen
30.53B (3.34B active) Partial offload 11/48 layers 18.64 GB 14.6 t/s
Qwen3 Next 80B A3B Instruct
Qwen
81.32B (3.19B active) Partial offload 4/48 layers 47.2 GB 13.3 t/s
Qwen3 Coder Next
Qwen
79.67B (3.19B active) Partial offload 4/48 layers 46.27 GB 13.3 t/s
GPT OSS 120B
GPT-OSS
116.83B (5.7B active) Partial offload 2/36 layers 66.48 GB 10.3 t/s
Phi 3.5 MoE Instruct
Phi
41.87B (6.64B active) Partial offload 5/32 layers 25.39 GB 7.6 t/s
Laguna S 2.1
Laguna
117.56B (7.71B active) Partial offload 3/48 layers 66.98 GB 7.5 t/s
DeepSeek Coder 7B Instruct V1.5
DeepSeek
6.91B Partial offload 17/30 layers 8.49 GB 7.4 t/s
DeepSeek Coder 6.7B Instruct
DeepSeek
6.74B Partial offload 17/32 layers 8.64 GB 6.9 t/s
CodeLlama 7B
Llama
6.74B Partial offload 17/32 layers 8.64 GB 6.9 t/s
Qwen1.5 7B
Qwen
7.72B Partial offload 16/32 layers 9.19 GB 6.2 t/s
Qwen3.5 122B A10B
Qwen
125.09B (8.17B active) Partial offload 2/48 layers 71.88 GB 6.1 t/s
Falcon 7B
Falcon
7.22B Partial offload 16/32 layers 9.38 GB 6.1 t/s
Qwen3 14B
Qwen
14.77B Partial offload 17/40 layers 10.47 GB 4.9 t/s
Phi 4
Phi
14.66B Partial offload 17/40 layers 10.72 GB 4.8 t/s

Showing the 90 best results of 133. The remaining 43 need more memory than this device has, at any quantisation.

What it will not run

9 of the 133 architectures we track are out of reach here, even at two-bit precision. Another 52 run by splitting layers between the card and system RAM, which works but drops generation to single digits.

If you want more headroom

The next steps up in memory, in order. More memory changes which models load at all; more bandwidth changes how fast they answer, so the two are worth weighing separately.

DeviceMemoryBandwidthModels that fit
GeForce RTX 3070 Ti 8 GB 608 GB/s 83 Check price
Arc A750 8 GB 512 GB/s 83 Check price
GeForce RTX 5060 Ti 8GB 8 GB 448 GB/s 83 Check price
GeForce RTX 5060 8 GB 448 GB/s 83 Check price

Price links go to an Amazon search for the model name and are affiliate links: if you buy through one, we earn a commission at no cost to you. We do not take payment for placement, and the ordering above is by memory capacity alone.

Other devices with 6 GB

Same capacity, different speed. Once a model fits, bandwidth is what separates these.