Unified memory · NVIDIA
What AI models can a Jetson Orin Nano Super 8GB run?
8 GB is the awkward size. It runs the 7 to 9 billion parameter class comfortably and hits a wall immediately above it, so most of the decisions here are about context length rather than model choice. Memory is shared with the CPU, so the practical ceiling is lower than the sticker number: macOS hands roughly three quarters of it to the GPU by default.
The short answer
Assuming an 8k context window and default settings, these are the models worth downloading first.
See how fast it feels
Jetson Orin Nano Super 8GB running LFM2.5 8B A1B at IQ4_XS
YouWhy does my model use more memory when the conversation gets longer?
Because of the KV cache. Every token you send leaves behind a key and a value vector in each layer of the model, and those stay in memory for as long as the conversation lasts.
The weights are a fixed cost: load a 4-bit 8B model and that is about 4.8 GB, whether you write one word or ten thousand. The cache is the part that grows, and it grows in a straight line with the number of tokens in the window.
How steeply depends on the model's attention design. With grouped-query attention, several query heads share one key-value pair, which cuts the cache by that ratio. Without it, every head keeps its own, and a long context can cost more memory than the weights themselves.
If you are short on memory, the first thing to try is lowering the context window in your runtime, not the quantisation.
Reading your question
Simulated from our estimate at a 512-token question, not a recording. Assumes nothing else is competing for the GPU.
Every model, scored on this device
Each row uses the highest-quality quantisation that both fits and stays conversational. Speed is a single-stream estimate at 8k context.
| Model | Params | Verdict | Download | Memory used | Speed | Max context |
|---|---|---|---|---|---|---|
| LFM2.5 8B A1B Liquid | 8.47B (1.57B active) | Just fits | IQ4_XS | 5.29 GB | 72.6 t/s | 8k |
| Gemma 4 E4B Instruct Gemma | 8B | Just fits | Q4_K_M | 5.3 GB | 18.4 t/s | 8k |
| Qwen2.5 VL 7B Instruct Qwen | 8.29B | Just fits | IQ4_XS | 5.36 GB | 18.4 t/s | 8k |
| Qwen2.5 7B Instruct Qwen | 7.62B | Just fits | IQ4_XS | 5.03 GB | 19.9 t/s | 8k |
| Gemma 4 E2B Instruct Gemma | 5.12B | Just fits | Q6_K | 4.62 GB | 21.3 t/s | 64k |
| OLMoE 1B 7B 0125 Instruct OLMo | 6.92B (1.28B active) | Just fits | IQ4_XS | 5.15 GB | 51.2 t/s | 4k |
| Internlm3 8B Instruct InternLM | 8.8B | Just fits | Q3_K_M | 5.23 GB | 19.1 t/s | 8k |
| Qwen3.5 4B Qwen | 4.66B | Just fits | Q6_K | 5.32 GB | 18.3 t/s | 8k |
| Agents A1 4B Other | 4.54B | Just fits | Q6_K | 5.23 GB | 18.7 t/s | 8k |
| Gemma 3 4B Instruct Gemma | 4.3B | Just fits | Q8_0 | 5.31 GB | 18.4 t/s | 8k |
| Qwen3 4B Qwen | 4.02B | Just fits | Q6_K | 4.95 GB | 19.9 t/s | 8k |
| Mistral 7B Instruct V0.3 Mistral | 7.25B | Just fits | Q3_K_M | 5.15 GB | 19.5 t/s | 8k |
| Mistral 7B Instruct V0.2 Mistral | 7.24B | Just fits | Q3_K_M | 5.15 GB | 19.5 t/s | 8k |
| Phi 4 Mini Instruct Phi | 3.84B | Just fits | Q6_K | 4.72 GB | 21.3 t/s | 8k |
| Phi 3 Mini 4k Instruct Phi | 3.82B | Just fits | Q8_0 | 5.32 GB | 18.5 t/s | 4k |
| Granite 4.1 3B Granite | 3.4B | Just fits | Q8_0 | 4.75 GB | 21 t/s | 8k |
| PowerMoE 3B PowerLM | 3.37B (0.88B active) | Just fits | Q8_0 | 4.53 GB | 61 t/s | 4k |
| Qwen3.5 9B Qwen | 9.65B | Just fits | IQ3_XXS ! | 5.29 GB | 18.8 t/s | 8k |
| Llama 3.2 3B Instruct Llama | 3.21B | Just fits | Q8_0 | 4.84 GB | 20.6 t/s | 8k |
| Granite 4.1 8B Granite | 8.79B | Just fits | IQ3_XXS ! | 5.23 GB | 19.1 t/s | 8k |
| Fanar 1 9B Instruct Fanar | 8.78B | Just fits | IQ3_XXS ! | 5.26 GB | 18.8 t/s | 4k |
| Qwen3 8B Qwen | 8.19B | Just fits | IQ3_XXS ! | 4.89 GB | 20.7 t/s | 8k |
| Granite 3.0 8B Instruct Granite | 8.17B | Just fits | IQ3_XXS ! | 5.01 GB | 20.1 t/s | 4k |
| T Lite Instruct 2.1 T-Lite | 8.19B | Just fits | IQ3_XXS ! | 4.89 GB | 20.7 t/s | 8k |
| Qwen2.5 3B Instruct Qwen | 3.09B | Runs great | Q8_0 | 4.07 GB | 25.1 t/s | 32k |
| SmolLM3 3B Base SmolLM | 3.08B | Runs well | Q8_0 | 4.34 GB | 23.2 t/s | 16k |
| Llama 3.1 8B Instruct Llama | 8.03B | Just fits | IQ3_XXS ! | 4.71 GB | 21.7 t/s | 8k |
| Apertus 8B Instruct 2509 Apertus | 8.05B | Just fits | IQ3_XXS ! | 4.72 GB | 21.6 t/s | 8k |
| Llama 3 Taiwan 8B Instruct Llama | 8.03B | Just fits | IQ3_XXS ! | 4.71 GB | 21.7 t/s | 8k |
| Starcoder2 3B StarCoder | 3.03B | Runs great | Q8_0 | 3.9 GB | 26.8 t/s | 16k |
| LFM2.5 2.6B Liquid | 2.7B | Runs great | Q8_0 | 3.87 GB | 26.6 t/s | 16k |
| Gemma 2 2B Instruct Gemma | 2.61B | Runs great | Q8_0 | 3.73 GB | 28 t/s | 8k |
| Phi 2 Phi | 2.78B | Just fits | Q5_K_M | 5.1 GB | 19.3 t/s | 2k |
| PowerLM 3B PowerLM | 3.51B | Just fits | IQ4_XS | 5.29 GB | 18.4 t/s | 4k |
| Qwen3.5 2B Qwen | 2.27B | Runs great | Q8_0 | 3.35 GB | 31.9 t/s | 32k |
| OneRec 1.7B Other | 2.13B | Runs great | Q8_0 | 3.71 GB | 28 t/s | 16k |
| Qwen3 1.7B Qwen | 2.03B | Runs great | Q8_0 | 3.61 GB | 29 t/s | 16k |
| DeepSeek Coder V2 Lite Instruct DeepSeek | 15.71B (2.74B active) | Just fits | IQ2_XXS ! | 4.73 GB | 93.5 t/s | 16k |
| vLLM Translategemma 12B Instruct Gemma | 13.19B | Just fits | IQ2_XXS ! | 4.37 GB | 23.6 t/s | 32k |
| DeepSeek R1 Distill Qwen 1.5B Qwen | 1.78B | Runs great | Q8_0 | 2.57 GB | 44.7 t/s | 128k |
| Phi 3 Vision 128k Instruct Phi | 4.15B | Just fits | IQ3_XXS ! | 5.27 GB | 18.7 t/s | 8k |
| Gemma 3 12B Instruct Gemma | 12.19B | Just fits | IQ2_XXS ! | 4.57 GB | 22.4 t/s | 16k |
| Mistral Nemo Instruct 2407 Mistral | 12.25B | Just fits | IQ2_XXS ! | 5.1 GB | 20 t/s | 8k |
| Gemma 4 12B Instruct Gemma | 11.96B | Runs great | IQ2_XXS ! | 4.08 GB | 25.8 t/s | 32k |
| Qwen3 1.7B Base Qwen | 1.72B | Runs great | Q8_0 | 3.3 GB | 32.5 t/s | 16k |
| SmolLM2 1.7B SmolLM | 1.71B | Runs great | Q8_0 | 3.92 GB | 26.2 t/s | 8k |
| Qwen2.5 1.5B Instruct Qwen | 1.54B | Runs great | Q8_0 | 2.44 GB | 48 t/s | 32k |
| Pythia 1.4B Pythia | 1.52B | Runs great | Q8_0 | 3.73 GB | 27.8 t/s | 2k |
| Gemma 2 9B Instruct Gemma | 9.24B | Runs well | IQ2_XXS ! | 4.35 GB | 23.7 t/s | 8k |
| OLMo 2 0425 1B OLMo | 1.48B | Runs great | Q8_0 | 3.19 GB | 33.9 t/s | 4k |
| OLMo 3 7B Instruct OLMo | 7.3B | Just fits | IQ2_XXS ! | 4.6 GB | 22.3 t/s | 32k |
| Llama 3.2 1B Instruct Llama | 1.24B | Runs great | Q8_0 | 2.2 GB | 56.6 t/s | 64k |
| LFM2.5 1.2B Instruct Liquid | 1.17B | Runs great | Q8_0 | 2.13 GB | 59.4 t/s | 64k |
| TinyLlama 1.1B Chat V1.0 Llama | 1.1B | Runs great | Q8_0 | 1.99 GB | 66.4 t/s | 2k |
| MiniCPM5 1B MiniCPM | 1.08B | Runs great | Q8_0 | 1.95 GB | 66.6 t/s | 64k |
| Gemma 3 1B Instruct Gemma | 1B | Runs great | Q8_0 | 1.71 GB | 80.5 t/s | 32k |
| Qwen3.5 0.8B Qwen | 0.87B | Runs great | Q8_0 | 1.9 GB | 67.7 t/s | 64k |
| Sarashina2.2 0.5B Instruct V0.1 Sarashina | 0.79B | Runs great | Q8_0 | 1.93 GB | 66.9 t/s | 8k |
| Qwen3 0.6B Qwen | 0.75B | Runs great | Q8_0 | 2.28 GB | 51.7 t/s | 32k |
| Qwen1.5 0.5B Chat Qwen | 0.62B | Runs great | Q8_0 | 2.03 GB | 61.3 t/s | 32k |
| Qwen3 0.6B Base Qwen | 0.6B | Runs great | Q8_0 | 2.13 GB | 56.9 t/s | 32k |
| Pythia 410m Pythia | 0.51B | Runs great | Q8_0 | 1.92 GB | 66.7 t/s | 2k |
| H2o Danube3 500m Chat Danube | 0.51B | Runs great | Q8_0 | 1.57 GB | 95.1 t/s | 8k |
| Qwen2.5 0.5B Instruct Qwen | 0.49B | Runs great | Q8_0 | 1.23 GB | 144.6 t/s | 32k |
| SmolLM2 360M SmolLM | 0.36B | Runs great | Q8_0 | 1.33 GB | 125.1 t/s | 8k |
| LFM2.5 350M Liquid | 0.35B | Runs great | Q8_0 | 1.26 GB | 140.3 t/s | 64k |
| Pythia 160m Pythia | 0.21B | Runs great | Q8_0 | 1.14 GB | 171 t/s | 2k |
| Japanese GPT NeoX Small GPT-NeoX | 0.2B | Runs great | Q8_0 | 1.13 GB | 174.6 t/s | 2k |
| Llama 160m Llama | 0.16B | Runs great | Q8_0 | 1.09 GB | 190.3 t/s | 2k |
| LLM Jp 3 150m LLM-jp | 0.15B | Runs great | Q8_0 | 0.97 GB | 249 t/s | 4k |
| SmolLM2 135M SmolLM | 0.13B | Runs great | Q8_0 | 0.94 GB | 274.8 t/s | 8k |
| Pythia 70m Deduped Pythia | 0.1B | Runs great | Q8_0 | 0.82 GB | 434 t/s | 2k |
| Mixtral 8x7B Instruct V0.1 Mistral | 46.7B (12.88B active) | CPU only | — | 28.11 GB | — | — |
| Phi 3.5 MoE Instruct Phi | 41.87B (6.64B active) | CPU only | — | 25.39 GB | — | — |
| Qwen3.5 35B A3B Qwen | 35.95B (2.9B active) | CPU only | — | 21.57 GB | — | — |
| Qwen3 32B Qwen | 32.76B | CPU only | — | 21.33 GB | — | — |
| Qwen2.5 32B Instruct Qwen | 32.76B | CPU only | — | 21.33 GB | — | — |
| OTel 2.0 LLM 31B Instruct Other | 32.11B | CPU only | — | 19.92 GB | — | — |
| Gemma 4 31B Instruct Gemma | 31.27B | CPU only | — | 19.45 GB | — | — |
| GLM 4.7 Flash GLM | 31.22B (3.66B active) | CPU only | — | 18.69 GB | — | — |
| Qwen3 30B A3B Qwen | 30.53B (3.34B active) | CPU only | — | 18.64 GB | — | — |
| Granite 4.1 30B Granite | 28.87B | CPU only | — | 19.08 GB | — | — |
| Qwen3.5 27B Qwen | 27.78B | CPU only | — | 18.53 GB | — | — |
| Gemma 3 27B Instruct Gemma | 27.43B | CPU only | — | 17.44 GB | — | — |
| Gemma 4 26B A4B Instruct Gemma | 25.81B | CPU only | — | 15.52 GB | — | — |
| Mistral Small 24B Instruct 2501 Mistral | 23.57B | CPU only | — | 15.42 GB | — | — |
| Codestral 22B V0.1 Mistral | 22.25B | CPU only | — | 15.24 GB | — | — |
| GPT OSS 20B GPT-OSS | 20.91B (4.18B active) | CPU only | — | 12.54 GB | — | — |
| GPT NeoX 20B GPT-NeoX | 20.74B | CPU only | — | 20.89 GB | — | — |
| Qwen3 14B Qwen | 14.77B | CPU only | — | 10.47 GB | — | — |
Showing the 90 best results of 133. The remaining 43 need more memory than this device has, at any quantisation.
What it will not run
61 of the 133 architectures we track are out of reach here, even at two-bit precision.
If you want more headroom
The next steps up in memory, in order. More memory changes which models load at all; more bandwidth changes how fast they answer, so the two are worth weighing separately.
| Device | Memory | Bandwidth | Models that fit | |
|---|---|---|---|---|
| GeForce RTX 3080 10GB | 10 GB | 760 GB/s | 88 | Check price |
| Arc B570 | 10 GB | 380 GB/s | 88 | Check price |
| GeForce RTX 2080 Ti | 11 GB | 616 GB/s | 93 | Check price |
| GeForce GTX 1080 Ti | 11 GB | 484 GB/s | 93 | Check price |
Price links go to an Amazon search for the model name and are affiliate links: if you buy through one, we earn a commission at no cost to you. We do not take payment for placement, and the ordering above is by memory capacity alone.
Other devices with 8 GB
Same capacity, different speed. Once a model fits, bandwidth is what separates these.