Save this page. It answers the only question that matters when you're buying for local AI: "what's the cheapest box that runs the model I want, and how fast?" Every number below is an owner-reported real-world speed (not a spec-sheet guess), aggregated from r/LocalLLaMA and cited benchmark sources. Pair it with our two tools when you're ready to pull the trigger: the quant picker (which exact file to download) and Can I run it? (what fits your machine).
Which box runs which model
| Model size | Cheapest box that runs it well | Real owner speed (Q4) |
|---|---|---|
| 8B (Llama 3.1 8B) | RTX 3060 12GB (~$289), almost any 12GB+ GPU | ~60–140 tok/s |
| 14B (Qwen 14B) | RTX 3060 12GB / any 16GB card | ~30–55 tok/s |
| 32B (Qwen 32B) | Used RTX 3090 24GB (~$700) | ~28–33 tok/s (3090) · ~45–60 (5090) |
👇 The 70B, 120B-MoE and 235B-MoE rows, plus five money-saving rules and the one-line buying rule, are free for subscribers. Sign up to see the rest.