Catalogue · updated 2026-09-01

Every model, and what it really costs in memory

133 architectures, 40 of them mixture-of-experts designs that occupy the memory of a large model while running at the speed of a small one. Every figure below is derived from the model's own configuration file rather than a rule of thumb, which is why two models of the same size can want very different amounts of memory.

Most downloaded

Weights shown at Q4_K_M, cache at an 8k window.

ModelFamilyParamsWeights at Q4Cache at 8kContextDownloads
Qwen3 0.6B Qwen 0.75B 0.4 GB 0.88 GB 40k 22.7M
Qwen3 8B Qwen 8.19B 4.6 GB 1.13 GB 40k 13.6M
Qwen3.5 9B Qwen 9.65B 5.4 GB 1.00 GB 256k 12.6M
Qwen2.5 7B Instruct Qwen 7.62B 4.3 GB 0.44 GB 32k 10.6M
Gemma 4 31B Instruct Gemma 31.27B 17.6 GB 0.94 GB 256k 8.3M
Gemma 4 26B A4B Instruct Gemma 25.81B 14.5 GB 0.23 GB 256k 8.2M
Qwen2.5 VL 7B Instruct Qwen 8.29B 4.7 GB 0.44 GB 125k 8.0M
Qwen2.5 1.5B Instruct Qwen 1.54B 0.9 GB 0.22 GB 32k 7.7M
Qwen2.5 3B Instruct Qwen 3.09B 1.7 GB 0.28 GB 32k 7.6M
Qwen3.5 4B Qwen 4.66B 2.6 GB 1.00 GB 256k 7.4M
OTel 2.0 LLM 31B Instruct Other 32.11B 18.1 GB 0.94 GB 256k 6.8M
Llama 3.2 1B Instruct Llama 1.24B 0.7 GB 0.25 GB 128k 6.7M

Under 3B

Runs on anything, including a phone or a laptop with no dedicated GPU. 33 models in this range.

and 15 more in this range.

3B to 9B

The sweet spot for an 8 GB card. Good general chat, weak at hard reasoning. 37 models in this range.

and 19 more in this range.

9B to 32B

Needs 12 to 24 GB. This is where local models start feeling genuinely useful. 23 models in this range.

and 5 more in this range.

32B to 80B

A 24 GB card at low precision, or unified memory. Slow but strong. 12 models in this range.

Over 80B

Multi-GPU, a big Mac, or a server. Mixture-of-experts models are the exception. 28 models in this range.

and 10 more in this range.

By family

GPT-OSS

2 models

MiMo

2 models

PowerLM

2 models

A.X

1 model

Apertus

1 model

Command

1 model

Danube

1 model

Falcon

1 model

Fanar

1 model

InternLM

1 model

Laguna

1 model

LLM-jp

1 model

MiniCPM

1 model

StarCoder

1 model

Step

1 model

T-Lite

1 model