Unified memory · AMD
What AI models can a Ryzen AI 9 HX 370 32GB run?
32 GB puts the 70B class within comfortable reach and leaves room for the context window that makes those models worth running. Memory is shared with the CPU, so the practical ceiling is lower than the sticker number: macOS hands roughly three quarters of it to the GPU by default.
The short answer
Assuming an 8k context window and default settings, these are the models worth downloading first.
Best all-rounder
GLM 4.7 Flash
Q5_K_M · 21.82 GB · about 29.3 tokens/s
Best for code
DeepSeek Coder 6.7B Instruct
Q4_K_M · 8.64 GB · about 10.7 tokens/s
Largest that still runs well
Phi 3.5 MoE Instruct
41.87B parameters · Q3_K_M · about 20.7 tokens/s
Fastest useful answer
PowerMoE 3B
about 60.7 tokens/s · 4.53 GB
See how fast it feels
Ryzen AI 9 HX 370 32GB running GLM 4.7 Flash at Q5_K_M
YouWhy does my model use more memory when the conversation gets longer?
Because of the KV cache. Every token you send leaves behind a key and a value vector in each layer of the model, and those stay in memory for as long as the conversation lasts.
The weights are a fixed cost: load a 4-bit 8B model and that is about 4.8 GB, whether you write one word or ten thousand. The cache is the part that grows, and it grows in a straight line with the number of tokens in the window.
How steeply depends on the model's attention design. With grouped-query attention, several query heads share one key-value pair, which cuts the cache by that ratio. Without it, every head keeps its own, and a long context can cost more memory than the weights themselves.
If you are short on memory, the first thing to try is lowering the context window in your runtime, not the quantisation.
Reading your question
Simulated from our estimate at a 512-token question, not a recording. Assumes nothing else is competing for the GPU.
Every model, scored on this device
Each row uses the highest-quality quantisation that both fits and stays conversational. Speed is a single-stream estimate at 8k context.
| Model | Params | Verdict | Download | Memory used | Speed | Max context |
|---|---|---|---|---|---|---|
| GLM 4.7 Flash GLM | 31.22B (3.66B active) | Just fits | Q5_K_M | 21.82 GB | 29.3 t/s | 16k |
| Qwen3 30B A3B Qwen | 30.53B (3.34B active) | Just fits | Q5_K_M | 21.7 GB | 28.1 t/s | 8k |
| Qwen3.5 35B A3B Qwen | 35.95B (2.9B active) | Just fits | Q4_K_M | 21.57 GB | 36.9 t/s | 16k |
| GPT OSS 20B GPT-OSS | 20.91B (4.18B active) | Just fits | Q8_0 | 21.47 GB | 20.1 t/s | 32k |
| Phi 3.5 MoE Instruct Phi | 41.87B (6.64B active) | Runs well | Q3_K_M | 20.91 GB | 20.7 t/s | 16k |
| DeepSeek Coder V2 Lite Instruct DeepSeek | 15.71B (2.74B active) | Runs great | Q8_0 | 16.51 GB | 28.2 t/s | 128k |
| Qwen1.5 MoE A2.7B Qwen | 14.32B (2.69B active) | Runs well | Q8_0 | 16.4 GB | 20 t/s | 8k |
| LFM2.5 8B A1B Liquid | 8.47B (1.57B active) | Runs great | Q8_0 | 9.48 GB | 43.1 t/s | 64k |
| Mixtral 8x7B Instruct V0.1 Mistral | 46.7B (12.88B active) | Runs well | IQ3_XXS ! | 18.49 GB | 14.9 t/s | 32k |
| DeepSeek V4 Flash 0731 Spark DeepSeek | 60.31B (20.36B active) | Just fits | IQ3_XXS ! | 22.35 GB | 11.5 t/s | 8k |
| OLMoE 1B 7B 0125 Instruct OLMo | 6.92B (1.28B active) | Runs great | Q8_0 | 8.57 GB | 36.7 t/s | 4k |
| Gemma 4 12B Instruct Gemma | 11.96B | Runs well | Q5_K_M | 9.13 GB | 10 t/s | 256k |
| vLLM Translategemma 12B Instruct Gemma | 13.19B | Runs well | Q4_K_M | 8.63 GB | 10.7 t/s | 128k |
| Internlm3 8B Instruct InternLM | 8.8B | Runs well | Q6_K | 7.95 GB | 11.7 t/s | 32k |
| Qwen2.5 VL 7B Instruct Qwen | 8.29B | Runs well | Q6_K | 7.59 GB | 12.3 t/s | 64k |
| Gemma 3 12B Instruct Gemma | 12.19B | Runs well | Q4_K_M | 8.5 GB | 10.9 t/s | 128k |
| Qwen3.5 9B Qwen | 9.65B | Runs well | Q5_K_M | 8.24 GB | 11.3 t/s | 64k |
| Llama 3.1 8B Instruct Llama | 8.03B | Runs well | Q6_K | 7.98 GB | 11.7 t/s | 64k |
| Mistral Nemo Instruct 2407 Mistral | 12.25B | Runs well | Q4_K_M | 9.05 GB | 10.2 t/s | 64k |
| Llama 3 Taiwan 8B Instruct Llama | 8.03B | Runs well | Q6_K | 7.98 GB | 11.7 t/s | 8k |
| Gemma 2 9B Instruct Gemma | 9.24B | Runs well | Q5_K_M | 8.25 GB | 11.2 t/s | 8k |
| Apertus 8B Instruct 2509 Apertus | 8.05B | Runs well | Q6_K | 8 GB | 11.6 t/s | 64k |
| Qwen3 8B Qwen | 8.19B | Runs well | Q6_K | 8.23 GB | 11.3 t/s | 32k |
| Qwen3 Next 80B A3B Instruct Qwen | 81.32B (3.19B active) | Runs great | IQ2_XXS ! | 20.98 GB | 54.9 t/s | 16k |
| T Lite Instruct 2.1 T-Lite | 8.19B | Runs well | Q6_K | 8.23 GB | 11.3 t/s | 32k |
| Gemma 4 26B A4B Instruct Gemma | 25.81B | Fits, but slow | Q6_K | 20.72 GB | 4.2 t/s | 64k |
| Granite 4.1 8B Granite | 8.79B | Runs well | Q6_K | 8.81 GB | 10.4 t/s | 64k |
| Qwen3 Coder Next Qwen | 79.67B (3.19B active) | Runs great | IQ2_XXS ! | 20.58 GB | 54.9 t/s | 16k |
| Fanar 1 9B Instruct Fanar | 8.78B | Runs well | Q6_K | 8.84 GB | 10.4 t/s | 4k |
| Granite 3.0 8B Instruct Granite | 8.17B | Runs well | Q6_K | 8.34 GB | 11.1 t/s | 4k |
| Gemma 3 27B Instruct Gemma | 27.43B | Fits, but slow | Q5_K_M | 20.19 GB | 4.3 t/s | 16k |
| Gemma 4 E4B Instruct Gemma | 8B | Runs well | Q8_0 | 8.72 GB | 10.5 t/s | 128k |
| Mistral Small 24B Instruct 2501 Mistral | 23.57B | Fits, but slow | Q6_K | 20.16 GB | 4.3 t/s | 16k |
| Gemma 4 31B Instruct Gemma | 31.27B | Fits, but slow | Q4_K_M | 19.45 GB | 4.5 t/s | 64k |
| OTel 2.0 LLM 31B Instruct Other | 32.11B | Fits, but slow | Q4_K_M | 19.92 GB | 4.4 t/s | 64k |
| Codestral 22B V0.1 Mistral | 22.25B | Fits, but slow | Q6_K | 19.72 GB | 4.4 t/s | 16k |
| Qwen3.5 27B Qwen | 27.78B | Fits, but slow | Q5_K_M | 21.32 GB | 4.1 t/s | 8k |
| Gemma 4 E2B Instruct Gemma | 5.12B | Runs well | Q8_0 | 5.78 GB | 16.4 t/s | 128k |
| Granite 4.1 30B Granite | 28.87B | Just fits | Q5_K_M | 21.97 GB | 3.9 t/s | 8k |
| Qwen2.5 7B Instruct Qwen | 7.62B | Runs well | Q8_0 | 8.8 GB | 10.4 t/s | 32k |
| Qwen3 32B Qwen | 32.76B | Fits, but slow | Q4_K_M | 21.33 GB | 4.1 t/s | 8k |
| Qwen2.5 32B Instruct Qwen | 32.76B | Fits, but slow | Q4_K_M | 21.33 GB | 4.1 t/s | 8k |
| OLMo 3 7B Instruct OLMo | 7.3B | Runs well | Q6_K | 8.43 GB | 11 t/s | 64k |
| Mistral 7B Instruct V0.3 Mistral | 7.25B | Runs well | Q8_0 | 9.02 GB | 10.2 t/s | 32k |
| Gemma 3 4B Instruct Gemma | 4.3B | Runs well | Q8_0 | 5.31 GB | 18.3 t/s | 128k |
| Mistral 7B Instruct V0.2 Mistral | 7.24B | Runs well | Q8_0 | 9.01 GB | 10.2 t/s | 32k |
| Qwen3 14B Qwen | 14.77B | Runs well | Q3_K_M | 8.89 GB | 10.4 t/s | 32k |
| Qwen2.5 14B Instruct Qwen | 14.77B | Runs well | Q3_K_M | 9.14 GB | 10.1 t/s | 32k |
| Phi 4 Phi | 14.66B | Runs well | Q3_K_M | 9.15 GB | 10.1 t/s | 16k |
| Qwen3.5 4B Qwen | 4.66B | Runs well | Q8_0 | 6.37 GB | 14.8 t/s | 64k |
| Agents A1 4B Other | 4.54B | Runs well | Q8_0 | 6.25 GB | 15.1 t/s | 64k |
| Phi 3 Mini 4k Instruct Phi | 3.82B | Runs well | Q8_0 | 5.32 GB | 18.4 t/s | 4k |
| Qwen3 4B Qwen | 4.02B | Runs well | Q8_0 | 5.86 GB | 16.3 t/s | 32k |
| Phi 4 Mini Instruct Phi | 3.84B | Runs well | Q8_0 | 5.59 GB | 17.3 t/s | 64k |
| Granite 4.1 3B Granite | 3.4B | Runs well | Q8_0 | 4.75 GB | 20.9 t/s | 128k |
| PowerMoE 3B PowerLM | 3.37B (0.88B active) | Runs great | Q8_0 | 4.53 GB | 60.7 t/s | 4k |
| DeepSeek Coder 7B Instruct V1.5 DeepSeek | 6.91B | Fits, but slow | Q5_K_M | 9.18 GB | 10 t/s | 4k |
| GPT NeoX 20B GPT-NeoX | 20.74B | Fits, but slow | Q4_K_M | 20.89 GB | 4.2 t/s | 2k |
| Qwen1.5 7B Qwen | 7.72B | Fits, but slow | Q4_K_M | 9.19 GB | 10 t/s | 32k |
| OLMo 2 1124 13B Instruct OLMo | 13.72B | Fits, but slow | Q8_0 | 20.74 GB | 4.2 t/s | 4k |
| Llama 3.2 3B Instruct Llama | 3.21B | Runs well | Q8_0 | 4.84 GB | 20.5 t/s | 128k |
| Qwen2.5 3B Instruct Qwen | 3.09B | Runs well | Q8_0 | 4.07 GB | 24.9 t/s | 32k |
| SmolLM3 3B Base SmolLM | 3.08B | Runs well | Q8_0 | 4.34 GB | 23 t/s | 64k |
| CodeLlama 7B Llama | 6.74B | Runs well | Q4_K_M | 8.64 GB | 10.7 t/s | 16k |
| DeepSeek Coder 6.7B Instruct DeepSeek | 6.74B | Runs well | Q4_K_M | 8.64 GB | 10.7 t/s | 16k |
| Starcoder2 3B StarCoder | 3.03B | Runs great | Q8_0 | 3.9 GB | 26.7 t/s | 16k |
| Falcon 7B Falcon | 7.22B | Runs well | IQ4_XS | 8.89 GB | 10.4 t/s | 8k |
| Phi 3 Vision 128k Instruct Phi | 4.15B | Runs well | Q8_0 | 7.89 GB | 11.7 t/s | 32k |
| LFM2.5 2.6B Liquid | 2.7B | Runs great | Q8_0 | 3.87 GB | 26.5 t/s | 128k |
| Gemma 2 2B Instruct Gemma | 2.61B | Runs great | Q8_0 | 3.73 GB | 27.8 t/s | 8k |
| PowerLM 3B PowerLM | 3.51B | Runs well | Q8_0 | 7.03 GB | 13.2 t/s | 4k |
| Phi 2 Phi | 2.78B | Runs well | Q8_0 | 6.01 GB | 15.8 t/s | 2k |
| Qwen3.5 2B Qwen | 2.27B | Runs great | Q8_0 | 3.35 GB | 31.7 t/s | 256k |
| OneRec 1.7B Other | 2.13B | Runs great | Q8_0 | 3.71 GB | 27.9 t/s | 32k |
| Qwen3 1.7B Qwen | 2.03B | Runs great | Q8_0 | 3.61 GB | 28.9 t/s | 32k |
| DeepSeek R1 Distill Qwen 1.5B Qwen | 1.78B | Runs great | Q8_0 | 2.57 GB | 44.5 t/s | 128k |
| Qwen3 1.7B Base Qwen | 1.72B | Runs great | Q8_0 | 3.3 GB | 32.3 t/s | 32k |
| SmolLM2 1.7B SmolLM | 1.71B | Runs great | Q8_0 | 3.92 GB | 26.1 t/s | 8k |
| Qwen2.5 1.5B Instruct Qwen | 1.54B | Runs great | Q8_0 | 2.44 GB | 47.7 t/s | 32k |
| Pythia 1.4B Pythia | 1.52B | Runs great | Q8_0 | 3.73 GB | 27.7 t/s | 2k |
| OLMo 2 0425 1B OLMo | 1.48B | Runs great | Q8_0 | 3.19 GB | 33.8 t/s | 4k |
| Llama 3.3 70B Instruct Llama | 70.55B | Fits, but slow | IQ2_XXS ! | 20.52 GB | 4.3 t/s | 8k |
| Qwen2.5 72B Instruct Qwen | 72.71B | Fits, but slow | IQ2_XXS ! | 21.04 GB | 4.2 t/s | 8k |
| Qwen2.5 VL 72B Instruct Qwen | 73.41B | Fits, but slow | IQ2_XXS ! | 21.21 GB | 4.1 t/s | 8k |
| Llama 3.2 1B Instruct Llama | 1.24B | Runs great | Q8_0 | 2.2 GB | 56.3 t/s | 128k |
| LFM2.5 1.2B Instruct Liquid | 1.17B | Runs great | Q8_0 | 2.13 GB | 59.1 t/s | 64k |
| TinyLlama 1.1B Chat V1.0 Llama | 1.1B | Runs great | Q8_0 | 1.99 GB | 66 t/s | 2k |
| MiniCPM5 1B MiniCPM | 1.08B | Runs great | Q8_0 | 1.95 GB | 66.2 t/s | 128k |
| Gemma 3 1B Instruct Gemma | 1B | Runs great | Q8_0 | 1.71 GB | 80.1 t/s | 32k |
| Qwen3.5 0.8B Qwen | 0.87B | Runs great | Q8_0 | 1.9 GB | 67.3 t/s | 256k |
Showing the 90 best results of 133. The remaining 43 need more memory than this device has, at any quantisation.
What it will not run
28 of the 133 architectures we track are out of reach here, even at two-bit precision.
If you want more headroom
The next steps up in memory, in order. More memory changes which models load at all; more bandwidth changes how fast they answer, so the two are worth weighing separately.
| Device | Memory | Bandwidth | Models that fit | |
|---|---|---|---|---|
| RTX 6000 Ada Generation | 48 GB | 960 GB/s | 111 | Check price |
| Radeon PRO W7900 | 48 GB | 864 GB/s | 111 | Check price |
| RTX A6000 | 48 GB | 768 GB/s | 111 | Check price |
| Ryzen AI Max+ 395 64GB | 64 GB | 256 GB/s | 111 | Check price |
Price links go to an Amazon search for the model name and are affiliate links: if you buy through one, we earn a commission at no cost to you. We do not take payment for placement, and the ordering above is by memory capacity alone.
Other devices with 32 GB
Same capacity, different speed. Once a model fits, bandwidth is what separates these.