Llama · 32B to 80B
Llama 3 3 Nemotron Super 49B V1
49.87 billion parameters across 80 layers. At four-bit precision the weights alone come to 28.04 GB, before any conversation is loaded. This one is cache-hungry: a 32k window adds 80.0 GB on top, so context length matters more here than the model size suggests.
Memory needed, by quantisation
At an 8k context window. Quality is our rough ranking of how much the compression costs you: anything at or above Q5 is hard to tell apart from the original in normal use.
| Quantisation | File size | KV cache | Total VRAM | Quality | What it costs you |
|---|---|---|---|---|---|
| Q8_0 | 49.35 GB | 20 GB | 70.45 GB | 99% | Lossless in practice. Use it when the memory is there. |
| Q6_K | 38.08 GB | 20 GB | 59.19 GB | 98% | Very close to Q8 for two thirds of the size. |
| Q5_K_M | 33.03 GB | 20 GB | 54.14 GB | 96% | The quality-first choice when Q6 will not fit. |
| Q4_K_M | 28.04 GB | 20 GB | 49.14 GB | 93% | The default. Best size-to-quality ratio for local use. |
| IQ4_XS | 24.67 GB | 20 GB | 45.77 GB | 91% | Importance-matrix 4-bit. Q4_K_S quality, smaller file. |
| Q3_K_M | 22.7 GB | 20 GB | 43.8 GB | 86% | Degradation starts to show. A way to fit one size up. |
| IQ3_XXS | 17.77 GB | 20 GB | 38.87 GB | 79% | Aggressive. Only worth it on very large models. |
What a longer conversation costs
Same model at Q4_K_M, only the context window changes. The cache grows in a straight line with every token in the window.
| Context | KV cache | Total VRAM | Fits in 8 GB | Fits in 12 GB | Fits in 24 GB |
|---|---|---|---|---|---|
| 4k | 10 GB | 38.89 GB | no | no | no |
| 8k | 20 GB | 49.14 GB | no | no | no |
| 16k | 40 GB | 69.64 GB | no | no | no |
| 32k | 80 GB | 110.64 GB | no | no | no |
| 64k | 160 GB | 192.64 GB | no | no | no |
| 128k | 320 GB | 356.64 GB | no | no | no |
The "fits" columns allow for the roughly 0.8 GB Windows keeps for the desktop.
Which hardware runs Llama 3 3 Nemotron Super 49B V1
26 of 118 consumer devices run it at a quantisation worth using, at an 8k context window. Another 3 can load it only by compressing the weights far enough to damage the model, marked with a warning below.
| Device | Memory | Verdict | Quantisation | Used | Speed |
|---|---|---|---|---|---|
| RTX PRO 6000 Blackwell NVIDIA | 96 GB | Runs well | Q8_0 | 70.45 GB | 21.2 t/s |
| RTX 6000 Ada Generation NVIDIA | 48 GB | Runs well | IQ4_XS | 45.77 GB | 17.6 t/s |
| Apple M3 Ultra 96GB Apple | 96 GB | Runs well | Q6_K | 58.89 GB | 11 t/s |
| Apple M3 Ultra 256GB Apple | 256 GB | Runs well | Q6_K | 58.89 GB | 11 t/s |
| Apple M3 Ultra 512GB Apple | 512 GB | Runs well | Q6_K | 58.89 GB | 11 t/s |
| Apple M1 Ultra 128GB Apple | 128 GB | Runs well | Q6_K | 58.89 GB | 10.7 t/s |
| Apple M2 Ultra 128GB Apple | 128 GB | Runs well | Q6_K | 58.89 GB | 10.7 t/s |
| Apple M2 Ultra 192GB Apple | 192 GB | Runs well | Q6_K | 58.89 GB | 10.7 t/s |
| RTX A6000 NVIDIA | 48 GB | Runs well | IQ4_XS | 45.77 GB | 14.1 t/s |
| Apple M1 Ultra 64GB Apple | 64 GB | Runs well | IQ4_XS | 45.47 GB | 14 t/s |
| Apple M2 Ultra 64GB Apple | 64 GB | Runs well | IQ4_XS | 45.47 GB | 14 t/s |
| Radeon PRO W7900 AMD | 48 GB | Runs well | IQ4_XS | 45.77 GB | 13.9 t/s |
| Apple M4 Max 64GB Apple | 64 GB | Fits, but slow | Q3_K_M | 43.5 GB | 10 t/s |
| Apple M4 Max 128GB Apple | 128 GB | Fits, but slow | Q3_K_M | 43.5 GB | 10 t/s |
| Apple M1 Max 64GB Apple | 64 GB | Fits, but slow | IQ4_XS | 45.47 GB | 7 t/s |
| Apple M2 Max 64GB Apple | 64 GB | Fits, but slow | IQ4_XS | 45.47 GB | 7 t/s |
| Apple M3 Max 64GB Apple | 64 GB | Fits, but slow | IQ4_XS | 45.47 GB | 7 t/s |
| Apple M2 Max 96GB Apple | 96 GB | Fits, but slow | Q8_0 | 70.15 GB | 4.5 t/s |
| Apple M3 Max 96GB Apple | 96 GB | Fits, but slow | Q8_0 | 70.15 GB | 4.5 t/s |
| Apple M3 Max 128GB Apple | 128 GB | Fits, but slow | Q8_0 | 70.15 GB | 4.5 t/s |
| NVIDIA DGX Spark 128GB NVIDIA | 128 GB | Fits, but slow | Q8_0 | 70.45 GB | 3.2 t/s |
| Apple M4 Pro 64GB Apple | 64 GB | Fits, but slow | IQ4_XS | 45.47 GB | 4.8 t/s |
| Jetson AGX Orin 64GB NVIDIA | 64 GB | Fits, but slow | IQ4_XS | 45.77 GB | 3.8 t/s |
| Ryzen AI Max+ 395 64GB AMD | 64 GB | Fits, but slow | IQ4_XS | 45.77 GB | 3.7 t/s |
| Ryzen AI Max+ 395 96GB AMD | 96 GB | Fits, but slow | Q8_0 | 70.45 GB | 2.4 t/s |
| Ryzen AI Max+ 395 128GB AMD | 128 GB | Fits, but slow | Q8_0 | 70.45 GB | 2.4 t/s |
| Apple M4 Max 48GB Apple | 48 GB | Runs well | IQ2_XXS ! | 32.76 GB | 13.3 t/s |
| Apple M3 Max 48GB Apple | 48 GB | Fits, but slow | IQ2_XXS ! | 32.76 GB | 9.8 t/s |
| Apple M4 Pro 48GB Apple | 48 GB | Fits, but slow | IQ2_XXS ! | 32.76 GB | 6.7 t/s |
| GeForce RTX 5090 NVIDIA | 32 GB | Partial offload 50/80 | — | 49.14 GB | 1.7 t/s |
| GeForce RTX 4090 NVIDIA | 24 GB | Partial offload 36/80 | — | 49.14 GB | 1.2 t/s |
| GeForce RTX 3090 Ti NVIDIA | 24 GB | Partial offload 36/80 | — | 49.14 GB | 1.2 t/s |
| GeForce RTX 3090 NVIDIA | 24 GB | Partial offload 36/80 | — | 49.14 GB | 1.2 t/s |
| GeForce RTX 5090 Laptop NVIDIA | 24 GB | Partial offload 36/80 | — | 49.14 GB | 1.2 t/s |
| Radeon RX 7900 XTX AMD | 24 GB | Partial offload 36/80 | — | 49.14 GB | 1.2 t/s |
| RTX A5000 NVIDIA | 24 GB | Partial offload 36/80 | — | 49.14 GB | 1.1 t/s |
| Radeon RX 7900 XT AMD | 20 GB | Partial offload 30/80 | — | 49.14 GB | 1 t/s |
| GeForce RTX 5080 NVIDIA | 16 GB | Partial offload 23/80 | — | 49.14 GB | 0.9 t/s |
| GeForce RTX 5070 Ti NVIDIA | 16 GB | Partial offload 23/80 | — | 49.14 GB | 0.9 t/s |
| GeForce RTX 5060 Ti 16GB NVIDIA | 16 GB | Partial offload 23/80 | — | 49.14 GB | 0.9 t/s |
| GeForce RTX 4080 SUPER NVIDIA | 16 GB | Partial offload 23/80 | — | 49.14 GB | 0.9 t/s |
| GeForce RTX 4080 NVIDIA | 16 GB | Partial offload 23/80 | — | 49.14 GB | 0.9 t/s |
| GeForce RTX 4070 Ti SUPER NVIDIA | 16 GB | Partial offload 23/80 | — | 49.14 GB | 0.9 t/s |
| GeForce RTX 4060 Ti 16GB NVIDIA | 16 GB | Partial offload 23/80 | — | 49.14 GB | 0.9 t/s |
| GeForce RTX 4090 Laptop NVIDIA | 16 GB | Partial offload 23/80 | — | 49.14 GB | 0.9 t/s |
| GeForce RTX 5080 Laptop NVIDIA | 16 GB | Partial offload 23/80 | — | 49.14 GB | 0.9 t/s |
| RTX A4000 NVIDIA | 16 GB | Partial offload 23/80 | — | 49.14 GB | 0.9 t/s |
| Radeon RX 9070 XT AMD | 16 GB | Partial offload 23/80 | — | 49.14 GB | 0.9 t/s |
| Radeon RX 9070 AMD | 16 GB | Partial offload 23/80 | — | 49.14 GB | 0.9 t/s |
| Radeon RX 7900 GRE AMD | 16 GB | Partial offload 23/80 | — | 49.14 GB | 0.9 t/s |
| Radeon RX 7800 XT AMD | 16 GB | Partial offload 23/80 | — | 49.14 GB | 0.9 t/s |
| Radeon RX 7600 XT AMD | 16 GB | Partial offload 23/80 | — | 49.14 GB | 0.9 t/s |
| Radeon RX 6900 XT AMD | 16 GB | Partial offload 23/80 | — | 49.14 GB | 0.9 t/s |
| Radeon RX 6800 AMD | 16 GB | Partial offload 23/80 | — | 49.14 GB | 0.9 t/s |
| Arc A770 16GB Intel | 16 GB | Partial offload 23/80 | — | 49.14 GB | 0.9 t/s |
| GeForce RTX 5070 NVIDIA | 12 GB | Partial offload 16/80 | — | 49.14 GB | 0.8 t/s |
| GeForce RTX 4070 Ti NVIDIA | 12 GB | Partial offload 16/80 | — | 49.14 GB | 0.8 t/s |
| GeForce RTX 4070 SUPER NVIDIA | 12 GB | Partial offload 16/80 | — | 49.14 GB | 0.8 t/s |
| GeForce RTX 4070 NVIDIA | 12 GB | Partial offload 16/80 | — | 49.14 GB | 0.8 t/s |
| GeForce RTX 3080 Ti NVIDIA | 12 GB | Partial offload 16/80 | — | 49.14 GB | 0.8 t/s |
What to buy to run Llama 3 3 Nemotron Super 49B V1
The cheapest hardware that runs it at a quantisation worth using and a speed you would not resent, at an 8k context window.
Cheapest that works
RTX 6000 Ada Generation
$6,800
IQ4_XS · uses 45.77 GB of its 48 GB · about 17.6 tokens/s
Best value
RTX PRO 6000 Blackwell
$8,500
Q8_0 · about 21.2 tokens/s · 1792 GB/s
If it just has to run
Apple M1 Max 64GB
$1,900 whole machine
IQ4_XS · about 7 tokens/s. Cheaper, and slower to answer.
| Device | Price | Memory | Runs it at | Speed | |
|---|---|---|---|---|---|
| RTX 6000 Ada Generation NVIDIA | $6,800 | 48 GB | IQ4_XS | 17.6 t/s | Check price |
| RTX PRO 6000 Blackwell NVIDIA | $8,500 | 96 GB | Q8_0 | 21.2 t/s | Check price |
Indicative prices reviewed 2026-09-01; used prices are marketplace typical. Options that cost more than a cheaper one with no more memory and no more speed are hidden. Price links are Amazon affiliate links. Change the standard, the context window or the budget in the buying tool.
Where it sits in the catalogue
Source: nvidia/Llama-3_3-Nemotron-Super-49B-v1 on Hugging Face. Downloaded 176k times in the last month. Published 2025-03-16. Architecture figures are read from the repository's own configuration file, so they move when the model does.
Llama 3.3 70B Instruct
70.55B · 128k context
Llama 3.1 8B Instruct
8.03B · 128k context
Llama 3 Taiwan 8B Instruct
8.03B · 8k context
CodeLlama 7B
6.74B · 16k context
Llama 3.2 3B Instruct
3.21B · 128k context
Llama 3.2 1B Instruct
1.24B · 128k context
TinyLlama 1.1B Chat V1.0
1.1B · 2k context
Llama 160m
0.16B · 2k context