Unified memory · Apple

What AI models can an Apple M1 Pro 32GB run?

32 GB puts the 70B class within comfortable reach and leaves room for the context window that makes those models worth running. Memory is shared with the CPU, so the practical ceiling is lower than the sticker number: macOS hands roughly three quarters of it to the GPU by default.

Memory32 GB
Bandwidth200 GB/s
Usable for a model22.4 GB
Runtime backendMETAL

The short answer

Assuming an 8k context window and default settings, these are the models worth downloading first.

See how fast it feels

Apple M1 Pro 32GB running GLM 4.7 Flash at Q5_K_M

Wait for the first word2.4 s
Then writes at55 tok/s
Whole answer5.5 s

YouWhy does my model use more memory when the conversation gets longer?

Model

Because of the KV cache. Every token you send leaves behind a key and a value vector in each layer of the model, and those stay in memory for as long as the conversation lasts.

The weights are a fixed cost: load a 4-bit 8B model and that is about 4.8 GB, whether you write one word or ten thousand. The cache is the part that grows, and it grows in a straight line with the number of tokens in the window.

How steeply depends on the model's attention design. With grouped-query attention, several query heads share one key-value pair, which cuts the cache by that ratio. Without it, every head keeps its own, and a long context can cost more memory than the weights themselves.

If you are short on memory, the first thing to try is lowering the context window in your runtime, not the quantisation.

Simulated from our estimate at a 512-token question, not a recording. Assumes nothing else is competing for the GPU.

Every model, scored on this device

Each row uses the highest-quality quantisation that both fits and stays conversational. Speed is a single-stream estimate at 8k context.

Model Params Verdict Download Memory used Speed Max context
GLM 4.7 Flash
GLM
31.22B (3.66B active) Just fits Q5_K_M 21.52 GB 55 t/s 16k
Qwen3 30B A3B
Qwen
30.53B (3.34B active) Runs great Q5_K_M 21.4 GB 52.7 t/s 16k
Qwen3.5 35B A3B
Qwen
35.95B (2.9B active) Runs great Q4_K_M 21.27 GB 69.2 t/s 16k
Phi 3.5 MoE Instruct
Phi
41.87B (6.64B active) Just fits IQ4_XS 22.27 GB 36.4 t/s 8k
GPT OSS 20B
GPT-OSS
20.91B (4.18B active) Runs great Q8_0 21.17 GB 37.7 t/s 32k
DeepSeek Coder V2 Lite Instruct
DeepSeek
15.71B (2.74B active) Runs great Q8_0 16.21 GB 52.9 t/s 128k
Qwen1.5 MoE A2.7B
Qwen
14.32B (2.69B active) Runs great Q8_0 16.1 GB 37.5 t/s 8k
DeepSeek V4 Flash 0731 Spark
DeepSeek
60.31B (20.36B active) Just fits IQ3_XXS ! 22.05 GB 21.5 t/s 16k
Mixtral 8x7B Instruct V0.1
Mistral
46.7B (12.88B active) Runs great IQ3_XXS ! 18.19 GB 27.9 t/s 32k
Gemma 4 26B A4B Instruct
Gemma
25.81B Runs well Q4_K_M 15.22 GB 10.6 t/s 256k
LFM2.5 8B A1B
Liquid
8.47B (1.57B active) Runs great Q8_0 9.18 GB 80.9 t/s 64k
Qwen3 14B
Qwen
14.77B Runs well Q6_K 13.14 GB 12.5 t/s 32k
Qwen2.5 14B Instruct
Qwen
14.77B Runs well Q6_K 13.39 GB 12.2 t/s 32k
Phi 4
Phi
14.66B Runs well Q6_K 13.37 GB 12.2 t/s 16k
Mistral Small 24B Instruct 2501
Mistral
23.57B Runs well Q4_K_M 15.12 GB 10.8 t/s 32k
Gemma 3 27B Instruct
Gemma
27.43B Runs well IQ4_XS 15.29 GB 10.6 t/s 64k
Gemma 4 E4B Instruct
Gemma
8B Runs well Q8_0 8.42 GB 19.6 t/s 128k
Codestral 22B V0.1
Mistral
22.25B Runs well Q4_K_M 14.94 GB 10.9 t/s 32k
Internlm3 8B Instruct
InternLM
8.8B Runs well Q8_0 9.63 GB 17.2 t/s 32k
Qwen2.5 VL 7B Instruct
Qwen
8.29B Runs well Q8_0 9.16 GB 18.1 t/s 64k
Gemma 4 12B Instruct
Gemma
11.96B Runs well Q8_0 12.75 GB 12.8 t/s 256k
Qwen2.5 7B Instruct
Qwen
7.62B Runs well Q8_0 8.5 GB 19.6 t/s 32k
vLLM Translategemma 12B Instruct
Gemma
13.19B Runs well Q8_0 13.96 GB 11.6 t/s 128k
Qwen3 32B
Qwen
32.76B Fits, but slow Q4_K_M 21.03 GB 7.6 t/s 8k
Qwen2.5 32B Instruct
Qwen
32.76B Fits, but slow Q4_K_M 21.03 GB 7.6 t/s 8k
Gemma 3 12B Instruct
Gemma
12.19B Runs well Q8_0 13.41 GB 12.1 t/s 64k
Qwen3.5 9B
Qwen
9.65B Runs well Q8_0 11.1 GB 14.8 t/s 64k
Mistral Nemo Instruct 2407
Mistral
12.25B Runs well Q8_0 13.99 GB 11.7 t/s 32k
Llama 3.1 8B Instruct
Llama
8.03B Runs well Q8_0 9.5 GB 17.4 t/s 64k
Apertus 8B Instruct 2509
Apertus
8.05B Runs well Q8_0 9.52 GB 17.4 t/s 64k
Llama 3 Taiwan 8B Instruct
Llama
8.03B Runs well Q8_0 9.5 GB 17.4 t/s 8k
Qwen3 8B
Qwen
8.19B Runs well Q8_0 9.78 GB 16.9 t/s 32k
Mistral 7B Instruct V0.3
Mistral
7.25B Runs well Q8_0 8.72 GB 19.1 t/s 32k
T Lite Instruct 2.1
T-Lite
8.19B Runs well Q8_0 9.78 GB 16.9 t/s 32k
Mistral 7B Instruct V0.2
Mistral
7.24B Runs well Q8_0 8.71 GB 19.1 t/s 32k
Granite 4.1 8B
Granite
8.79B Runs well Q8_0 10.5 GB 15.7 t/s 64k
Gemma 2 9B Instruct
Gemma
9.24B Runs well Q8_0 10.98 GB 14.9 t/s 8k
OLMoE 1B 7B 0125 Instruct
OLMo
6.92B (1.28B active) Runs great Q8_0 8.27 GB 68.8 t/s 4k
Fanar 1 9B Instruct
Fanar
8.78B Runs well Q8_0 10.52 GB 15.6 t/s 4k
Granite 3.0 8B Instruct
Granite
8.17B Runs well Q8_0 9.88 GB 16.7 t/s 4k
Gemma 4 31B Instruct
Gemma
31.27B Runs well Q3_K_M 15.8 GB 10.3 t/s 128k
OTel 2.0 LLM 31B Instruct
Other
32.11B Runs well Q3_K_M 16.18 GB 10 t/s 128k
OLMo 3 7B Instruct
OLMo
7.3B Runs well Q8_0 9.77 GB 16.9 t/s 64k
Qwen3.5 27B
Qwen
27.78B Runs well Q3_K_M 15.26 GB 10.7 t/s 32k
Granite 4.1 30B
Granite
28.87B Runs well Q3_K_M 15.69 GB 10.3 t/s 16k
OLMo 2 1124 13B Instruct
OLMo
13.72B Runs well Q5_K_M 15.95 GB 10.2 t/s 4k
GPT NeoX 20B
GPT-NeoX
20.74B Fits, but slow Q4_K_M 20.59 GB 7.8 t/s 2k
Qwen1.5 7B
Qwen
7.72B Runs well Q8_0 12.19 GB 13.4 t/s 16k
DeepSeek Coder 7B Instruct V1.5
DeepSeek
6.91B Runs well Q8_0 11.14 GB 14.7 t/s 4k
Gemma 4 E2B Instruct
Gemma
5.12B Runs great Q8_0 5.48 GB 30.7 t/s 128k
CodeLlama 7B
Llama
6.74B Runs well Q8_0 11.22 GB 14.6 t/s 16k
DeepSeek Coder 6.7B Instruct
DeepSeek
6.74B Runs well Q8_0 11.22 GB 14.6 t/s 16k
Falcon 7B
Falcon
7.22B Runs well Q8_0 12.16 GB 13.5 t/s 8k
Qwen3.5 4B
Qwen
4.66B Runs great Q8_0 6.07 GB 27.8 t/s 64k
Qwen3 Next 80B A3B Instruct
Qwen
81.32B (3.19B active) Runs great IQ2_XXS ! 20.68 GB 103 t/s 16k
Qwen3 Coder Next
Qwen
79.67B (3.19B active) Runs great IQ2_XXS ! 20.28 GB 103 t/s 16k
Agents A1 4B
Other
4.54B Runs great Q8_0 5.95 GB 28.4 t/s 64k
Gemma 3 4B Instruct
Gemma
4.3B Runs great Q8_0 5.01 GB 34.3 t/s 128k
Phi 3 Vision 128k Instruct
Phi
4.15B Runs well Q8_0 7.59 GB 22 t/s 32k
Qwen3 4B
Qwen
4.02B Runs great Q8_0 5.56 GB 30.6 t/s 32k
Phi 4 Mini Instruct
Phi
3.84B Runs great Q8_0 5.29 GB 32.5 t/s 64k
Phi 3 Mini 4k Instruct
Phi
3.82B Runs great Q8_0 5.02 GB 34.4 t/s 4k
PowerLM 3B
PowerLM
3.51B Runs well Q8_0 6.73 GB 24.8 t/s 4k
Granite 4.1 3B
Granite
3.4B Runs great Q8_0 4.45 GB 39.1 t/s 128k
PowerMoE 3B
PowerLM
3.37B (0.88B active) Runs great Q8_0 4.23 GB 113.8 t/s 4k
Llama 3.2 3B Instruct
Llama
3.21B Runs great Q8_0 4.54 GB 38.5 t/s 128k
Qwen2.5 3B Instruct
Qwen
3.09B Runs great Q8_0 3.77 GB 46.7 t/s 32k
SmolLM3 3B Base
SmolLM
3.08B Runs great Q8_0 4.04 GB 43.2 t/s 64k
Starcoder2 3B
StarCoder
3.03B Runs great Q8_0 3.6 GB 50.1 t/s 16k
Phi 2
Phi
2.78B Runs great Q8_0 5.71 GB 29.7 t/s 2k
LFM2.5 2.6B
Liquid
2.7B Runs great Q8_0 3.57 GB 49.7 t/s 128k
Gemma 2 2B Instruct
Gemma
2.61B Runs great Q8_0 3.43 GB 52.2 t/s 8k
Qwen3.5 2B
Qwen
2.27B Runs great Q8_0 3.05 GB 59.5 t/s 256k
Llama 3.3 70B Instruct
Llama
70.55B Fits, but slow IQ2_XXS ! 20.22 GB 8 t/s 8k
Qwen2.5 VL 72B Instruct
Qwen
73.41B Fits, but slow IQ2_XXS ! 20.91 GB 7.8 t/s 8k
Qwen2.5 72B Instruct
Qwen
72.71B Fits, but slow IQ2_XXS ! 20.74 GB 7.8 t/s 8k
OneRec 1.7B
Other
2.13B Runs great Q8_0 3.41 GB 52.3 t/s 32k
Qwen3 1.7B
Qwen
2.03B Runs great Q8_0 3.31 GB 54.1 t/s 32k
DeepSeek R1 Distill Qwen 1.5B
Qwen
1.78B Runs great Q8_0 2.27 GB 83.4 t/s 128k
Qwen3 1.7B Base
Qwen
1.72B Runs great Q8_0 3 GB 60.5 t/s 32k
SmolLM2 1.7B
SmolLM
1.71B Runs great Q8_0 3.62 GB 48.9 t/s 8k
Qwen2.5 1.5B Instruct
Qwen
1.54B Runs great Q8_0 2.14 GB 89.5 t/s 32k
Pythia 1.4B
Pythia
1.52B Runs great Q8_0 3.43 GB 51.9 t/s 2k
OLMo 2 0425 1B
OLMo
1.48B Runs great Q8_0 2.89 GB 63.3 t/s 4k
Llama 3.2 1B Instruct
Llama
1.24B Runs great Q8_0 1.9 GB 105.6 t/s 128k
LFM2.5 1.2B Instruct
Liquid
1.17B Runs great Q8_0 1.83 GB 110.8 t/s 64k
TinyLlama 1.1B Chat V1.0
Llama
1.1B Runs great Q8_0 1.69 GB 123.8 t/s 2k
MiniCPM5 1B
MiniCPM
1.08B Runs great Q8_0 1.65 GB 124.2 t/s 128k
Gemma 3 1B Instruct
Gemma
1B Runs great Q8_0 1.41 GB 150.2 t/s 32k
Qwen3.5 0.8B
Qwen
0.87B Runs great Q8_0 1.6 GB 126.2 t/s 256k

Showing the 90 best results of 133. The remaining 43 need more memory than this device has, at any quantisation.

What it will not run

28 of the 133 architectures we track are out of reach here, even at two-bit precision.

If you want more headroom

The next steps up in memory, in order. More memory changes which models load at all; more bandwidth changes how fast they answer, so the two are worth weighing separately.

DeviceMemoryBandwidthModels that fit
RTX 6000 Ada Generation 48 GB 960 GB/s 111 Check price
Radeon PRO W7900 48 GB 864 GB/s 111 Check price
RTX A6000 48 GB 768 GB/s 111 Check price
Ryzen AI Max+ 395 64GB 64 GB 256 GB/s 111 Check price

Price links go to an Amazon search for the model name and are affiliate links: if you buy through one, we earn a commission at no cost to you. We do not take payment for placement, and the ordering above is by memory capacity alone.

Other devices with 32 GB

Same capacity, different speed. Once a model fits, bandwidth is what separates these.