Guides
The parts worth understanding
The calculator gives you a number. These explain where the number comes from, which is what you need when you are deciding what to buy rather than what to download.
How much VRAM do I need for a local AI model?
The honest answer is a formula, not a number. Here is the formula, what each term does, and the sizes it works out to for the cards people actually own.
Which GPU should you buy for local AI?
Memory decides what loads, bandwidth decides how fast it answers, and the two are separate purchases. How to pick, and why the obvious card is often the wrong one.
Quantisation explained: which one to download
What Q4_K_M and IQ3_XXS actually mean, what each level costs in quality, and the rule for choosing between a big model squeezed small and a small one kept intact.
Why a long context window eats your VRAM
The KV cache is the memory nobody budgets for, and why a model that loaded this morning will not load now. What it is, and the three designs that keep it small.