Step · Over 80B
Step 3.5 Flash
199.38 billion parameters across 45 layers. At four-bit precision the weights alone come to 112.11 GB, before any conversation is loaded. Grouped-query attention keeps the cache small, so long contexts cost less here than on models of the same size.
Memory needed, by quantisation
At an 8k context window. Quality is our rough ranking of how much the compression costs you: anything at or above Q5 is hard to tell apart from the original in normal use.
| Quantisation | File size | KV cache | Total VRAM | Quality | What it costs you |
|---|---|---|---|---|---|
| Q8_0 | 197.29 GB | 0.7 GB | 198.85 GB | 99% | Lossless in practice. Use it when the memory is there. |
| Q6_K | 152.26 GB | 0.7 GB | 153.82 GB | 98% | Very close to Q8 for two thirds of the size. |
| Q5_K_M | 132.07 GB | 0.7 GB | 133.62 GB | 96% | The quality-first choice when Q6 will not fit. |
| Q4_K_M | 112.11 GB | 0.7 GB | 113.66 GB | 93% | The default. Best size-to-quality ratio for local use. |
| IQ4_XS | 98.65 GB | 0.7 GB | 100.2 GB | 91% | Importance-matrix 4-bit. Q4_K_S quality, smaller file. |
| Q3_K_M | 90.75 GB | 0.7 GB | 92.31 GB | 86% | Degradation starts to show. A way to fit one size up. |
| IQ3_XXS | 71.03 GB | 0.7 GB | 72.58 GB | 79% | Aggressive. Only worth it on very large models. |
What a longer conversation costs
Same model at Q4_K_M, only the context window changes. This model uses sliding-window attention, so most layers stop growing past 512 tokens and the bill flattens out.
| Context | KV cache | Total VRAM | Fits in 8 GB | Fits in 12 GB | Fits in 24 GB |
|---|---|---|---|---|---|
| 4k | 0.7 GB | 113.54 GB | no | no | no |
| 8k | 0.7 GB | 113.66 GB | no | no | no |
| 16k | 0.7 GB | 113.91 GB | no | no | no |
| 32k | 0.7 GB | 114.41 GB | no | no | no |
| 64k | 0.7 GB | 115.41 GB | no | no | no |
| 128k | 0.7 GB | 117.41 GB | no | no | no |
The "fits" columns allow for the roughly 0.8 GB Windows keeps for the desktop.
Which hardware runs Step 3.5 Flash
10 of 118 consumer devices run it at a quantisation worth using, at an 8k context window. Another 4 can load it only by compressing the weights far enough to damage the model, marked with a warning below.
| Device | Memory | Verdict | Quantisation | Used | Speed |
|---|---|---|---|---|---|
| RTX PRO 6000 Blackwell NVIDIA | 96 GB | Runs well | Q3_K_M | 92.31 GB | 16.1 t/s |
| Apple M2 Ultra 192GB Apple | 192 GB | Fits, but slow | Q5_K_M | 133.32 GB | 4.7 t/s |
| Apple M3 Ultra 256GB Apple | 256 GB | Fits, but slow | Q6_K | 153.52 GB | 4.2 t/s |
| Apple M1 Ultra 128GB Apple | 128 GB | Fits, but slow | Q3_K_M | 92.01 GB | 6.8 t/s |
| Apple M2 Ultra 128GB Apple | 128 GB | Fits, but slow | Q3_K_M | 92.01 GB | 6.8 t/s |
| Apple M3 Ultra 512GB Apple | 512 GB | Fits, but slow | Q8_0 | 198.55 GB | 3.2 t/s |
| Apple M4 Max 128GB Apple | 128 GB | Fits, but slow | Q3_K_M | 92.01 GB | 4.7 t/s |
| Apple M3 Max 128GB Apple | 128 GB | Fits, but slow | Q3_K_M | 92.01 GB | 3.4 t/s |
| Apple M3 Ultra 96GB Apple | 96 GB | Runs well | IQ2_XXS ! | 49.07 GB | 13.2 t/s |
| NVIDIA DGX Spark 128GB NVIDIA | 128 GB | Fits, but slow | Q3_K_M | 92.31 GB | 2.4 t/s |
| Ryzen AI Max+ 395 128GB AMD | 128 GB | Fits, but slow | Q3_K_M | 92.31 GB | 1.8 t/s |
| Apple M2 Max 96GB Apple | 96 GB | Fits, but slow | IQ2_XXS ! | 49.07 GB | 6.4 t/s |
| Apple M3 Max 96GB Apple | 96 GB | Fits, but slow | IQ2_XXS ! | 49.07 GB | 6.4 t/s |
| Ryzen AI Max+ 395 96GB AMD | 96 GB | Fits, but slow | IQ2_XXS ! | 49.37 GB | 3.4 t/s |
| RTX 6000 Ada Generation NVIDIA | 48 GB | Partial offload 18/45 | — | 113.66 GB | 0.5 t/s |
| RTX A6000 NVIDIA | 48 GB | Partial offload 18/45 | — | 113.66 GB | 0.5 t/s |
| Radeon PRO W7900 AMD | 48 GB | Partial offload 18/45 | — | 113.66 GB | 0.5 t/s |
| GeForce RTX 5090 NVIDIA | 32 GB | Partial offload 12/45 | — | 113.66 GB | 0.4 t/s |
| GeForce RTX 5080 NVIDIA | 16 GB | Partial offload 5/45 | — | 113.66 GB | 0.3 t/s |
| GeForce RTX 5070 Ti NVIDIA | 16 GB | Partial offload 5/45 | — | 113.66 GB | 0.3 t/s |
| GeForce RTX 5070 NVIDIA | 12 GB | Partial offload 4/45 | — | 113.66 GB | 0.3 t/s |
| GeForce RTX 5060 Ti 16GB NVIDIA | 16 GB | Partial offload 5/45 | — | 113.66 GB | 0.3 t/s |
| GeForce RTX 5060 Ti 8GB NVIDIA | 8 GB | Partial offload 2/45 | — | 113.66 GB | 0.3 t/s |
| GeForce RTX 5060 NVIDIA | 8 GB | Partial offload 2/45 | — | 113.66 GB | 0.3 t/s |
| GeForce RTX 4090 NVIDIA | 24 GB | Partial offload 8/45 | — | 113.66 GB | 0.3 t/s |
| GeForce RTX 4080 SUPER NVIDIA | 16 GB | Partial offload 5/45 | — | 113.66 GB | 0.3 t/s |
| GeForce RTX 4080 NVIDIA | 16 GB | Partial offload 5/45 | — | 113.66 GB | 0.3 t/s |
| GeForce RTX 4070 Ti SUPER NVIDIA | 16 GB | Partial offload 5/45 | — | 113.66 GB | 0.3 t/s |
| GeForce RTX 4070 Ti NVIDIA | 12 GB | Partial offload 4/45 | — | 113.66 GB | 0.3 t/s |
| GeForce RTX 4070 SUPER NVIDIA | 12 GB | Partial offload 4/45 | — | 113.66 GB | 0.3 t/s |
| GeForce RTX 4070 NVIDIA | 12 GB | Partial offload 4/45 | — | 113.66 GB | 0.3 t/s |
| GeForce RTX 4060 Ti 16GB NVIDIA | 16 GB | Partial offload 5/45 | — | 113.66 GB | 0.3 t/s |
| GeForce RTX 4060 Ti 8GB NVIDIA | 8 GB | Partial offload 2/45 | — | 113.66 GB | 0.3 t/s |
| GeForce RTX 4060 NVIDIA | 8 GB | Partial offload 2/45 | — | 113.66 GB | 0.3 t/s |
| GeForce RTX 3090 Ti NVIDIA | 24 GB | Partial offload 8/45 | — | 113.66 GB | 0.3 t/s |
| GeForce RTX 3090 NVIDIA | 24 GB | Partial offload 8/45 | — | 113.66 GB | 0.3 t/s |
| GeForce RTX 3080 Ti NVIDIA | 12 GB | Partial offload 4/45 | — | 113.66 GB | 0.3 t/s |
| GeForce RTX 3080 12GB NVIDIA | 12 GB | Partial offload 4/45 | — | 113.66 GB | 0.3 t/s |
| GeForce RTX 3080 10GB NVIDIA | 10 GB | Partial offload 3/45 | — | 113.66 GB | 0.3 t/s |
| GeForce RTX 3070 Ti NVIDIA | 8 GB | Partial offload 2/45 | — | 113.66 GB | 0.3 t/s |
| GeForce RTX 3070 NVIDIA | 8 GB | Partial offload 2/45 | — | 113.66 GB | 0.3 t/s |
| GeForce RTX 3060 Ti NVIDIA | 8 GB | Partial offload 2/45 | — | 113.66 GB | 0.3 t/s |
| GeForce RTX 3060 12GB NVIDIA | 12 GB | Partial offload 4/45 | — | 113.66 GB | 0.3 t/s |
| GeForce RTX 3050 8GB NVIDIA | 8 GB | Partial offload 2/45 | — | 113.66 GB | 0.3 t/s |
| GeForce RTX 3050 6GB NVIDIA | 6 GB | Partial offload 1/45 | — | 113.66 GB | 0.3 t/s |
| GeForce RTX 2080 Ti NVIDIA | 11 GB | Partial offload 3/45 | — | 113.66 GB | 0.3 t/s |
| GeForce RTX 2060 12GB NVIDIA | 12 GB | Partial offload 4/45 | — | 113.66 GB | 0.3 t/s |
| GeForce GTX 1080 Ti NVIDIA | 11 GB | Partial offload 3/45 | — | 113.66 GB | 0.3 t/s |
| GeForce GTX 1660 SUPER NVIDIA | 6 GB | Partial offload 1/45 | — | 113.66 GB | 0.3 t/s |
| GeForce RTX 4090 Laptop NVIDIA | 16 GB | Partial offload 5/45 | — | 113.66 GB | 0.3 t/s |
| GeForce RTX 4080 Laptop NVIDIA | 12 GB | Partial offload 4/45 | — | 113.66 GB | 0.3 t/s |
| GeForce RTX 4070 Laptop NVIDIA | 8 GB | Partial offload 2/45 | — | 113.66 GB | 0.3 t/s |
| GeForce RTX 4060 Laptop NVIDIA | 8 GB | Partial offload 2/45 | — | 113.66 GB | 0.3 t/s |
| GeForce RTX 5090 Laptop NVIDIA | 24 GB | Partial offload 8/45 | — | 113.66 GB | 0.3 t/s |
| GeForce RTX 5080 Laptop NVIDIA | 16 GB | Partial offload 5/45 | — | 113.66 GB | 0.3 t/s |
| RTX A5000 NVIDIA | 24 GB | Partial offload 8/45 | — | 113.66 GB | 0.3 t/s |
| RTX A4000 NVIDIA | 16 GB | Partial offload 5/45 | — | 113.66 GB | 0.3 t/s |
| Radeon RX 9070 XT AMD | 16 GB | Partial offload 5/45 | — | 113.66 GB | 0.3 t/s |
| Radeon RX 9070 AMD | 16 GB | Partial offload 5/45 | — | 113.66 GB | 0.3 t/s |
| Radeon RX 7900 XTX AMD | 24 GB | Partial offload 8/45 | — | 113.66 GB | 0.3 t/s |
What to buy to run Step 3.5 Flash
Nothing we track runs this model comfortably at 8k context, so there is no honest buying answer here.
At 8k context this model needs more memory than anything in our consumer catalogue can give it at a quantisation we would recommend. The realistic options are a shorter context window, a smaller model from the same family below, or multiple cards, which we do not model.
Where it sits in the catalogue
Source: stepfun-ai/Step-3.5-Flash on Hugging Face. Downloaded 155k times in the last month. Published 2026-02-01. Architecture figures are read from the repository's own configuration file, so they move when the model does.