GPUs for Local LLM

The Local-LLM Hardware Cheat-Sheet: Which Box Runs Which Model

A one-page cheat-sheet: the cheapest hardware that runs each model size (8B to 235B), with real owner-reported speeds, free for subscribers.

The Local-LLM Hardware Cheat-Sheet: Which Box Runs Which Model

Save this page. It answers the only question that matters when you're buying for local AI: "what's the cheapest box that runs the model I want, and how fast?" Every number below is an owner-reported real-world speed (not a spec-sheet guess), aggregated from r/LocalLLaMA and cited benchmark sources. Pair it with our two tools when you're ready to pull the trigger: the quant picker (which exact file to download) and Can I run it? (what fits your machine).

Which box runs which model

Model sizeCheapest box that runs it wellReal owner speed (Q4)
8B (Llama 3.1 8B)RTX 3060 12GB (~$289), almost any 12GB+ GPU~60–140 tok/s
14B (Qwen 14B)RTX 3060 12GB / any 16GB card~30–55 tok/s
32B (Qwen 32B)Used RTX 3090 24GB (~$700)~28–33 tok/s (3090) · ~45–60 (5090)

👇 The 70B, 120B-MoE and 235B-MoE rows, plus five money-saving rules and the one-line buying rule, are free for subscribers. Sign up to see the rest.

Get the Vetted Consumer newsletter

Reviews, buying advice, and field notes. Delivered monthly.

Almost there, check your inbox and click the confirmation link. ✓

Something went wrong, please try again, or email [email protected].