Data centre GPU · AMD
What AI models can an Instinct MI300X run?
192 GB is server territory. The constraint stops being memory and becomes bandwidth: everything fits, and what varies is how long you wait for it. At 5300 GB/s there is enough bandwidth to keep generation responsive for anything that fits.
The short answer
Assuming an 8k context window and default settings, these are the models worth downloading first.
Best all-rounder
Qwen3 235B A22B
Q6_K · 181.85 GB · about 207.7 tokens/s
Best for code
Laguna S 2.1
Q8_0 · 117.21 GB · about 494.1 tokens/s
Largest that still runs well
Ornith 1.5 397B
403.4B parameters · Q3_K_M · about 502.6 tokens/s
Fastest useful answer
PowerMoE 3B
about 2783.8 tokens/s · 4.53 GB
See how fast it feels
Instinct MI300X running Qwen3 235B A22B at Q6_K
YouWhy does my model use more memory when the conversation gets longer?
Because of the KV cache. Every token you send leaves behind a key and a value vector in each layer of the model, and those stay in memory for as long as the conversation lasts.
The weights are a fixed cost: load a 4-bit 8B model and that is about 4.8 GB, whether you write one word or ten thousand. The cache is the part that grows, and it grows in a straight line with the number of tokens in the window.
How steeply depends on the model's attention design. With grouped-query attention, several query heads share one key-value pair, which cuts the cache by that ratio. Without it, every head keeps its own, and a long context can cost more memory than the weights themselves.
If you are short on memory, the first thing to try is lowering the context window in your runtime, not the quantisation.
Reading your question
Simulated from our estimate at a 512-token question, not a recording. Assumes nothing else is competing for the GPU.
Every model, scored on this device
Each row uses the highest-quality quantisation that both fits and stays conversational. Speed is a single-stream estimate at 8k context.
| Model | Params | Verdict | Download | Memory used | Speed | Max context |
|---|---|---|---|---|---|---|
| Qwen3 235B A22B Qwen | 235.09B (22.14B active) | Runs great | Q6_K | 181.85 GB | 207.7 t/s | 32k |
| MiniMax M2.7 MiniMax | 228.69B (10.98B active) | Runs great | Q6_K | 177.37 GB | 369.7 t/s | 32k |
| MiniMax M2.5 MiniMax | 228.7B (10.98B active) | Runs great | Q6_K | 177.38 GB | 369.7 t/s | 32k |
| Command A Plus 05 2026 Bf16 Command | 218.75B (18.52B active) | Runs great | Q6_K | 168.41 GB | 260.6 t/s | 128k |
| Step 3.5 Flash Step | 199.38B | Runs well | Q6_K | 153.82 GB | 24.9 t/s | 256k |
| DeepSeek V4 Flash DSpark DeepSeek | 165.27B (20.36B active) | Runs great | Q8_0 | 164.4 GB | 189.3 t/s | 256k |
| MiMo V2.5 MiMo | 310.78B (16.05B active) | Runs great | Q4_K_M | 175.62 GB | 422 t/s | 256k |
| MiMo V2 Flash MiMo | 309.79B (16.05B active) | Runs great | Q4_K_M | 175.06 GB | 422 t/s | 256k |
| DeepSeek V4 Flash 0731 DeepSeek | 304.18B (20.36B active) | Runs great | Q4_K_M | 171.9 GB | 333 t/s | 256k |
| DeepSeek V4 Flash DeepSeek | 290.94B (20.36B active) | Runs great | Q4_K_M | 164.45 GB | 333 t/s | 256k |
| GLM 4.5 GLM | 358.34B (33.56B active) | Runs great | IQ4_XS | 181.08 GB | 195.9 t/s | 32k |
| Qwen3.5 122B A10B Qwen | 125.09B (8.17B active) | Runs great | Q8_0 | 125.32 GB | 431.9 t/s | 256k |
| Laguna S 2.1 Laguna | 117.56B (7.71B active) | Runs great | Q8_0 | 117.21 GB | 494.1 t/s | 1M |
| GPT OSS 120B GPT-OSS | 116.83B (5.7B active) | Runs great | Q8_0 | 116.39 GB | 675.5 t/s | 128k |
| GLM 4.5 Air GLM | 110.47B (13.4B active) | Runs great | Q8_0 | 111.6 GB | 259.6 t/s | 128k |
| Ornith 1.5 397B Ornith | 403.4B (14.62B active) | Runs great | Q3_K_M | 185.41 GB | 502.6 t/s | 32k |
| Ornith 1.0 397B Ornith | 396.8B (14.62B active) | Runs great | Q3_K_M | 182.41 GB | 502.6 t/s | 64k |
| Qwen3 Next 80B A3B Instruct Qwen | 81.32B (3.19B active) | Runs great | Q8_0 | 81.94 GB | 976.8 t/s | 256k |
| Qwen3 Coder Next Qwen | 79.67B (3.19B active) | Runs great | Q8_0 | 80.31 GB | 976.8 t/s | 256k |
| Qwen2.5 VL 72B Instruct Qwen | 73.41B | Runs great | Q8_0 | 76.24 GB | 50.8 t/s | 64k |
| Qwen2.5 72B Instruct Qwen | 72.71B | Runs great | Q8_0 | 75.55 GB | 51.3 t/s | 32k |
| Llama 3.3 70B Instruct Llama | 70.55B | Runs great | Q8_0 | 73.41 GB | 52.8 t/s | 128k |
| DeepSeek V4 Flash 0731 Spark DeepSeek | 60.31B (20.36B active) | Runs great | Q8_0 | 60.54 GB | 189.3 t/s | 1M |
| Llama 3 3 Nemotron Super 49B V1 Llama | 49.87B | Runs great | Q8_0 | 70.45 GB | 55 t/s | 32k |
| Mixtral 8x7B Instruct V0.1 Mistral | 46.7B (12.88B active) | Runs great | Q8_0 | 48.06 GB | 277.6 t/s | 32k |
| MiniMax M3 MiniMax | 427.04B (25.86B active) | Runs great | IQ3_XXS ! | 154.04 GB | 376 t/s | 128k |
| Phi 3.5 MoE Instruct Phi | 41.87B (6.64B active) | Runs great | Q8_0 | 43.28 GB | 504.1 t/s | 128k |
| Qwen3.5 35B A3B Qwen | 35.95B (2.9B active) | Runs great | Q8_0 | 36.93 GB | 1092 t/s | 256k |
| Qwen3 32B Qwen | 32.76B | Runs great | Q8_0 | 35.33 GB | 110.9 t/s | 32k |
| Qwen2.5 32B Instruct Qwen | 32.76B | Runs great | Q8_0 | 35.33 GB | 110.9 t/s | 32k |
| OTel 2.0 LLM 31B Instruct Other | 32.11B | Runs great | Q8_0 | 33.64 GB | 116.7 t/s | 256k |
| Gemma 4 31B Instruct Gemma | 31.27B | Runs great | Q8_0 | 32.81 GB | 119.7 t/s | 256k |
| GLM 4.7 Flash GLM | 31.22B (3.66B active) | Runs great | Q8_0 | 32.03 GB | 945.8 t/s | 128k |
| Qwen3 30B A3B Qwen | 30.53B (3.34B active) | Runs great | Q8_0 | 31.69 GB | 941.1 t/s | 32k |
| Granite 4.1 30B Granite | 28.87B | Runs great | Q8_0 | 31.42 GB | 124.8 t/s | 128k |
| Qwen3.5 27B Qwen | 27.78B | Runs great | Q8_0 | 30.4 GB | 129.4 t/s | 256k |
| Gemma 3 27B Instruct Gemma | 27.43B | Runs great | Q8_0 | 29.16 GB | 135.2 t/s | 128k |
| Gemma 4 26B A4B Instruct Gemma | 25.81B | Runs great | Q8_0 | 26.55 GB | 148.1 t/s | 256k |
| Mistral Small 24B Instruct 2501 Mistral | 23.57B | Runs great | Q8_0 | 25.49 GB | 155.3 t/s | 32k |
| Codestral 22B V0.1 Mistral | 22.25B | Runs great | Q8_0 | 24.74 GB | 160.6 t/s | 32k |
| GPT OSS 20B GPT-OSS | 20.91B (4.18B active) | Runs great | Q8_0 | 21.47 GB | 921.3 t/s | 128k |
| GPT NeoX 20B GPT-NeoX | 20.74B | Runs great | Q8_0 | 29.75 GB | 132.6 t/s | 2k |
| DeepSeek Coder V2 Lite Instruct DeepSeek | 15.71B (2.74B active) | Runs great | Q8_0 | 16.51 GB | 1294.2 t/s | 128k |
| Qwen2.5 14B Instruct Qwen | 14.77B | Runs great | Q8_0 | 17.03 GB | 236.8 t/s | 32k |
| Qwen3 14B Qwen | 14.77B | Runs great | Q8_0 | 16.78 GB | 240.5 t/s | 32k |
| Phi 4 Phi | 14.66B | Runs great | Q8_0 | 16.98 GB | 237.5 t/s | 16k |
| Qwen1.5 MoE A2.7B Qwen | 14.32B (2.69B active) | Runs great | Q8_0 | 16.4 GB | 916.9 t/s | 8k |
| OLMo 2 1124 13B Instruct OLMo | 13.72B | Runs great | Q8_0 | 20.74 GB | 192.5 t/s | 4k |
| vLLM Translategemma 12B Instruct Gemma | 13.19B | Runs great | Q8_0 | 14.26 GB | 284.2 t/s | 128k |
| GLM 5.2 GLM | 753.33B (51.62B active) | Runs great | IQ2_XXS ! | 182.32 GB | 292.1 t/s | 64k |
| GLM 5.1 GLM | 753.86B (35.91B active) | Runs great | IQ2_XXS ! | 182.45 GB | 410.4 t/s | 64k |
| Mistral Nemo Instruct 2407 Mistral | 12.25B | Runs great | Q8_0 | 14.29 GB | 285.4 t/s | 128k |
| Gemma 3 12B Instruct Gemma | 12.19B | Runs great | Q8_0 | 13.71 GB | 296.4 t/s | 128k |
| A.X K2 A.X | 691.69B (39.06B active) | Runs great | IQ2_XXS ! | 167.45 GB | 385.3 t/s | 128k |
| Gemma 4 12B Instruct Gemma | 11.96B | Runs great | Q8_0 | 13.05 GB | 312.5 t/s | 256k |
| DeepSeek R1 DeepSeek | 684.53B (38.57B active) | Runs great | IQ2_XXS ! | 165.74 GB | 390 t/s | 128k |
| DeepSeek V3.2 DeepSeek | 685.4B (38.57B active) | Runs great | IQ2_XXS ! | 165.94 GB | 390 t/s | 128k |
| Qwen3.5 9B Qwen | 9.65B | Runs great | Q8_0 | 11.4 GB | 361.7 t/s | 256k |
| Gemma 2 9B Instruct Gemma | 9.24B | Runs great | Q8_0 | 11.28 GB | 365 t/s | 8k |
| Granite 4.1 8B Granite | 8.79B | Runs great | Q8_0 | 10.8 GB | 383.6 t/s | 128k |
| Fanar 1 9B Instruct Fanar | 8.78B | Runs great | Q8_0 | 10.82 GB | 381.6 t/s | 4k |
| Internlm3 8B Instruct InternLM | 8.8B | Runs great | Q8_0 | 9.93 GB | 420.1 t/s | 32k |
| LFM2.5 8B A1B Liquid | 8.47B (1.57B active) | Runs great | Q8_0 | 9.48 GB | 1978.7 t/s | 64k |
| Qwen2.5 VL 7B Instruct Qwen | 8.29B | Runs great | Q8_0 | 9.46 GB | 441.6 t/s | 64k |
| Qwen3 8B Qwen | 8.19B | Runs great | Q8_0 | 10.08 GB | 413.5 t/s | 32k |
| Granite 3.0 8B Instruct Granite | 8.17B | Runs great | Q8_0 | 10.18 GB | 408.8 t/s | 4k |
| T Lite Instruct 2.1 T-Lite | 8.19B | Runs great | Q8_0 | 10.08 GB | 413.5 t/s | 32k |
| Llama 3.1 8B Instruct Llama | 8.03B | Runs great | Q8_0 | 9.8 GB | 426.6 t/s | 128k |
| Apertus 8B Instruct 2509 Apertus | 8.05B | Runs great | Q8_0 | 9.82 GB | 425.6 t/s | 64k |
| Llama 3 Taiwan 8B Instruct Llama | 8.03B | Runs great | Q8_0 | 9.8 GB | 426.6 t/s | 8k |
| Gemma 4 E4B Instruct Gemma | 8B | Runs great | Q8_0 | 8.72 GB | 479.6 t/s | 128k |
| Qwen1.5 7B Qwen | 7.72B | Runs great | Q8_0 | 12.49 GB | 327.9 t/s | 32k |
| Qwen2.5 7B Instruct Qwen | 7.62B | Runs great | Q8_0 | 8.8 GB | 478.3 t/s | 32k |
| OLMo 3 7B Instruct OLMo | 7.3B | Runs great | Q8_0 | 10.07 GB | 413.7 t/s | 64k |
| Mistral 7B Instruct V0.3 Mistral | 7.25B | Runs great | Q8_0 | 9.02 GB | 466.8 t/s | 32k |
| Mistral 7B Instruct V0.2 Mistral | 7.24B | Runs great | Q8_0 | 9.01 GB | 467.4 t/s | 32k |
| Falcon 7B Falcon | 7.22B | Runs great | Q8_0 | 12.46 GB | 329.5 t/s | 8k |
| DeepSeek Coder 7B Instruct V1.5 DeepSeek | 6.91B | Runs great | Q8_0 | 11.44 GB | 360.4 t/s | 4k |
| OLMoE 1B 7B 0125 Instruct OLMo | 6.92B (1.28B active) | Runs great | Q8_0 | 8.57 GB | 1683.6 t/s | 4k |
| CodeLlama 7B Llama | 6.74B | Runs great | Q8_0 | 11.52 GB | 357.7 t/s | 16k |
| DeepSeek Coder 6.7B Instruct DeepSeek | 6.74B | Runs great | Q8_0 | 11.52 GB | 357.7 t/s | 16k |
| Gemma 4 E2B Instruct Gemma | 5.12B | Runs great | Q8_0 | 5.78 GB | 750.7 t/s | 128k |
| Qwen3.5 4B Qwen | 4.66B | Runs great | Q8_0 | 6.37 GB | 680.1 t/s | 256k |
| Agents A1 4B Other | 4.54B | Runs great | Q8_0 | 6.25 GB | 694.8 t/s | 256k |
| Gemma 3 4B Instruct Gemma | 4.3B | Runs great | Q8_0 | 5.31 GB | 838.3 t/s | 128k |
| Phi 3 Vision 128k Instruct Phi | 4.15B | Runs great | Q8_0 | 7.89 GB | 537 t/s | 128k |
| Qwen3 4B Qwen | 4.02B | Runs great | Q8_0 | 5.86 GB | 747.8 t/s | 32k |
| Phi 4 Mini Instruct Phi | 3.84B | Runs great | Q8_0 | 5.59 GB | 795 t/s | 128k |
| Phi 3 Mini 4k Instruct Phi | 3.82B | Runs great | Q8_0 | 5.32 GB | 842.5 t/s | 4k |
| PowerLM 3B PowerLM | 3.51B | Runs great | Q8_0 | 7.03 GB | 607.1 t/s | 4k |
Showing the 90 best results of 133. The remaining 43 need more memory than this device has, at any quantisation.
What it will not run
0 of the 133 architectures we track are out of reach here, even at two-bit precision. Another 4 run by splitting layers between the card and system RAM, which works but drops generation to single digits.
Other devices with 192 GB
Same capacity, different speed. Once a model fits, bandwidth is what separates these.