Unified memory · Apple

What AI models can an Apple M3 Ultra 512GB run?

512 GB is server territory. The constraint stops being memory and becomes bandwidth: everything fits, and what varies is how long you wait for it. Memory is shared with the CPU, so the practical ceiling is lower than the sticker number: macOS hands roughly three quarters of it to the GPU by default.

Memory512 GB
Bandwidth819 GB/s
Usable for a model384 GB
Runtime backendMETAL

The short answer

Assuming an 8k context window and default settings, these are the models worth downloading first.

See how fast it feels

Apple M3 Ultra 512GB running MiniMax M3 at Q6_K

Wait for the first word3.2 s
Then writes at30.9 tok/s
Whole answer8.8 s

YouWhy does my model use more memory when the conversation gets longer?

Model

Because of the KV cache. Every token you send leaves behind a key and a value vector in each layer of the model, and those stay in memory for as long as the conversation lasts.

The weights are a fixed cost: load a 4-bit 8B model and that is about 4.8 GB, whether you write one word or ten thousand. The cache is the part that grows, and it grows in a straight line with the number of tokens in the window.

How steeply depends on the model's attention design. With grouped-query attention, several query heads share one key-value pair, which cuts the cache by that ratio. Without it, every head keeps its own, and a long context can cost more memory than the weights themselves.

If you are short on memory, the first thing to try is lowering the context window in your runtime, not the quantisation.

Simulated from our estimate at a 512-token question, not a recording. Assumes nothing else is competing for the GPU.

Every model, scored on this device

Each row uses the highest-quality quantisation that both fits and stays conversational. Speed is a single-stream estimate at 8k context.

Model Params Verdict Download Memory used Speed Max context
MiniMax M3
MiniMax
427.04B (25.86B active) Runs great Q6_K 327.74 GB 30.9 t/s 256k
Ornith 1.5 397B
Ornith
403.4B (14.62B active) Runs great Q6_K 309.56 GB 52.8 t/s 256k
Ornith 1.0 397B
Ornith
396.8B (14.62B active) Runs great Q6_K 304.52 GB 52.8 t/s 256k
MiMo V2.5
MiMo
310.78B (16.05B active) Runs great Q8_0 308.09 GB 40.2 t/s 1M
MiMo V2 Flash
MiMo
309.79B (16.05B active) Runs great Q8_0 307.12 GB 40.2 t/s 256k
DeepSeek V4 Flash 0731
DeepSeek
304.18B (20.36B active) Runs great Q8_0 301.56 GB 31.7 t/s 1M
DeepSeek V4 Flash
DeepSeek
290.94B (20.36B active) Runs great Q8_0 288.46 GB 31.7 t/s 1M
GLM 4.5
GLM
358.34B (33.56B active) Runs well Q8_0 358.08 GB 17.7 t/s 64k
GLM 5.2
GLM
753.33B (51.62B active) Runs well IQ4_XS 374.08 GB 24.4 t/s 64k
GLM 5.1
GLM
753.86B (35.91B active) Runs great IQ4_XS 374.35 GB 34.6 t/s 64k
A.X K2
A.X
691.69B (39.06B active) Runs great IQ4_XS 343.5 GB 32.2 t/s 256k
DeepSeek R1
DeepSeek
684.53B (38.57B active) Runs great IQ4_XS 339.96 GB 32.6 t/s 128k
DeepSeek V3.2
DeepSeek
685.4B (38.57B active) Runs great IQ4_XS 340.39 GB 32.6 t/s 128k
Qwen3 235B A22B
Qwen
235.09B (22.14B active) Runs great Q8_0 234.65 GB 27.3 t/s 32k
MiniMax M2.7
MiniMax
228.69B (10.98B active) Runs great Q8_0 228.72 GB 49.9 t/s 128k
MiniMax M2.5
MiniMax
228.7B (10.98B active) Runs great Q8_0 228.73 GB 49.9 t/s 128k
Command A Plus 05 2026 Bf16
Command
218.75B (18.52B active) Runs great Q8_0 217.51 GB 33.9 t/s 128k
DeepSeek V4 Flash DSpark
DeepSeek
165.27B (20.36B active) Runs great Q8_0 164.1 GB 31.7 t/s 1M
Qwen3.5 122B A10B
Qwen
125.09B (8.17B active) Runs great Q8_0 125.02 GB 72.3 t/s 256k
Laguna S 2.1
Laguna
117.56B (7.71B active) Runs great Q8_0 116.91 GB 82.7 t/s 1M
GPT OSS 120B
GPT-OSS
116.83B (5.7B active) Runs great Q8_0 116.09 GB 113.1 t/s 128k
GLM 4.5 Air
GLM
110.47B (13.4B active) Runs great Q8_0 111.3 GB 43.5 t/s 128k
Kimi K2.6
Kimi
1026.88B (39.06B active) Runs great IQ3_XXS ! 367.08 GB 44.2 t/s 128k
Kimi K2 Instruct
Kimi
1026.41B (39.06B active) Runs great IQ3_XXS ! 366.91 GB 44.2 t/s 128k
Qwen3 Next 80B A3B Instruct
Qwen
81.32B (3.19B active) Runs great Q8_0 81.64 GB 163.5 t/s 256k
Qwen3 Coder Next
Qwen
79.67B (3.19B active) Runs great Q8_0 80.01 GB 163.5 t/s 256k
DeepSeek V4 Flash 0731 Spark
DeepSeek
60.31B (20.36B active) Runs great Q8_0 60.24 GB 31.7 t/s 1M
Mixtral 8x7B Instruct V0.1
Mistral
46.7B (12.88B active) Runs great Q8_0 47.76 GB 46.5 t/s 32k
Phi 3.5 MoE Instruct
Phi
41.87B (6.64B active) Runs great Q8_0 42.98 GB 84.4 t/s 128k
Qwen3.5 35B A3B
Qwen
35.95B (2.9B active) Runs great Q8_0 36.63 GB 182.8 t/s 256k
Gemma 4 31B Instruct
Gemma
31.27B Runs well Q8_0 32.51 GB 20 t/s 256k
GLM 4.7 Flash
GLM
31.22B (3.66B active) Runs great Q8_0 31.73 GB 158.3 t/s 128k
OTel 2.0 LLM 31B Instruct
Other
32.11B Runs well Q8_0 33.34 GB 19.5 t/s 256k
Qwen3 30B A3B
Qwen
30.53B (3.34B active) Runs great Q8_0 31.39 GB 157.5 t/s 32k
Qwen3 32B
Qwen
32.76B Runs well Q8_0 35.03 GB 18.6 t/s 32k
Qwen2.5 32B Instruct
Qwen
32.76B Runs well Q8_0 35.03 GB 18.6 t/s 32k
Granite 4.1 30B
Granite
28.87B Runs well Q8_0 31.12 GB 20.9 t/s 128k
Qwen3.5 27B
Qwen
27.78B Runs well Q8_0 30.1 GB 21.7 t/s 256k
Gemma 3 27B Instruct
Gemma
27.43B Runs well Q8_0 28.86 GB 22.6 t/s 128k
Llama 3.3 70B Instruct
Llama
70.55B Runs well Q6_K 57.18 GB 11.3 t/s 128k
Qwen2.5 72B Instruct
Qwen
72.71B Runs well Q6_K 58.83 GB 11 t/s 32k
Qwen2.5 VL 72B Instruct
Qwen
73.41B Runs well Q6_K 59.36 GB 10.9 t/s 64k
Gemma 4 26B A4B Instruct
Gemma
25.81B Runs well Q8_0 26.25 GB 24.8 t/s 256k
Mistral Small 24B Instruct 2501
Mistral
23.57B Runs great Q8_0 25.19 GB 26 t/s 32k
Codestral 22B V0.1
Mistral
22.25B Runs great Q8_0 24.44 GB 26.9 t/s 32k
GPT OSS 20B
GPT-OSS
20.91B (4.18B active) Runs great Q8_0 21.17 GB 154.2 t/s 128k
GPT NeoX 20B
GPT-NeoX
20.74B Runs well Q8_0 29.45 GB 22.2 t/s 2k
Llama 3 3 Nemotron Super 49B V1
Llama
49.87B Runs well Q6_K 58.89 GB 11 t/s 128k
DeepSeek Coder V2 Lite Instruct
DeepSeek
15.71B (2.74B active) Runs great Q8_0 16.21 GB 216.7 t/s 128k
Qwen2.5 14B Instruct
Qwen
14.77B Runs great Q8_0 16.73 GB 39.6 t/s 32k
Qwen3 14B
Qwen
14.77B Runs great Q8_0 16.48 GB 40.3 t/s 32k
Phi 4
Phi
14.66B Runs great Q8_0 16.68 GB 39.8 t/s 16k
Qwen1.5 MoE A2.7B
Qwen
14.32B (2.69B active) Runs great Q8_0 16.1 GB 153.5 t/s 8k
OLMo 2 1124 13B Instruct
OLMo
13.72B Runs great Q8_0 20.44 GB 32.2 t/s 4k
vLLM Translategemma 12B Instruct
Gemma
13.19B Runs great Q8_0 13.96 GB 47.6 t/s 128k
Mistral Nemo Instruct 2407
Mistral
12.25B Runs great Q8_0 13.99 GB 47.8 t/s 128k
Gemma 3 12B Instruct
Gemma
12.19B Runs great Q8_0 13.41 GB 49.6 t/s 128k
Gemma 4 12B Instruct
Gemma
11.96B Runs great Q8_0 12.75 GB 52.3 t/s 256k
Step 3.5 Flash
Step
199.38B Fits, but slow Q8_0 198.55 GB 3.2 t/s 256k
Qwen3.5 9B
Qwen
9.65B Runs great Q8_0 11.1 GB 60.6 t/s 256k
Gemma 2 9B Instruct
Gemma
9.24B Runs great Q8_0 10.98 GB 61.1 t/s 8k
Granite 4.1 8B
Granite
8.79B Runs great Q8_0 10.5 GB 64.2 t/s 128k
Fanar 1 9B Instruct
Fanar
8.78B Runs great Q8_0 10.52 GB 63.9 t/s 4k
Internlm3 8B Instruct
InternLM
8.8B Runs great Q8_0 9.63 GB 70.3 t/s 32k
LFM2.5 8B A1B
Liquid
8.47B (1.57B active) Runs great Q8_0 9.18 GB 331.2 t/s 64k
Qwen2.5 VL 7B Instruct
Qwen
8.29B Runs great Q8_0 9.16 GB 73.9 t/s 64k
Qwen3 8B
Qwen
8.19B Runs great Q8_0 9.78 GB 69.2 t/s 32k
Granite 3.0 8B Instruct
Granite
8.17B Runs great Q8_0 9.88 GB 68.4 t/s 4k
T Lite Instruct 2.1
T-Lite
8.19B Runs great Q8_0 9.78 GB 69.2 t/s 32k
Llama 3.1 8B Instruct
Llama
8.03B Runs great Q8_0 9.5 GB 71.4 t/s 128k
Apertus 8B Instruct 2509
Apertus
8.05B Runs great Q8_0 9.52 GB 71.3 t/s 64k
Llama 3 Taiwan 8B Instruct
Llama
8.03B Runs great Q8_0 9.5 GB 71.4 t/s 8k
Gemma 4 E4B Instruct
Gemma
8B Runs great Q8_0 8.42 GB 80.3 t/s 128k
Qwen1.5 7B
Qwen
7.72B Runs great Q8_0 12.19 GB 54.9 t/s 32k
Qwen2.5 7B Instruct
Qwen
7.62B Runs great Q8_0 8.5 GB 80.1 t/s 32k
OLMo 3 7B Instruct
OLMo
7.3B Runs great Q8_0 9.77 GB 69.3 t/s 64k
Mistral 7B Instruct V0.3
Mistral
7.25B Runs great Q8_0 8.72 GB 78.2 t/s 32k
Mistral 7B Instruct V0.2
Mistral
7.24B Runs great Q8_0 8.71 GB 78.2 t/s 32k
Falcon 7B
Falcon
7.22B Runs great Q8_0 12.16 GB 55.2 t/s 8k
DeepSeek Coder 7B Instruct V1.5
DeepSeek
6.91B Runs great Q8_0 11.14 GB 60.3 t/s 4k
OLMoE 1B 7B 0125 Instruct
OLMo
6.92B (1.28B active) Runs great Q8_0 8.27 GB 281.8 t/s 4k
CodeLlama 7B
Llama
6.74B Runs great Q8_0 11.22 GB 59.9 t/s 16k
DeepSeek Coder 6.7B Instruct
DeepSeek
6.74B Runs great Q8_0 11.22 GB 59.9 t/s 16k
Gemma 4 E2B Instruct
Gemma
5.12B Runs great Q8_0 5.48 GB 125.7 t/s 128k
Qwen3.5 4B
Qwen
4.66B Runs great Q8_0 6.07 GB 113.8 t/s 256k
Agents A1 4B
Other
4.54B Runs great Q8_0 5.95 GB 116.3 t/s 256k
Gemma 3 4B Instruct
Gemma
4.3B Runs great Q8_0 5.01 GB 140.3 t/s 128k
Phi 3 Vision 128k Instruct
Phi
4.15B Runs great Q8_0 7.59 GB 89.9 t/s 128k
Qwen3 4B
Qwen
4.02B Runs great Q8_0 5.56 GB 125.2 t/s 32k
Phi 4 Mini Instruct
Phi
3.84B Runs great Q8_0 5.29 GB 133.1 t/s 128k

Showing the 90 best results of 133. The remaining 43 need more memory than this device has, at any quantisation.

What it will not run

2 of the 133 architectures we track are out of reach here, even at two-bit precision.