Category

Software & Tools

Honest guides to the software that runs local LLMs — inference engines, GUIs, quantization, serving, and RAG tooling. Mostly free, open-source tools, covered because they matter to the niche, not for affiliate links.

What "Open Weights" Lets You Do: The 2026 Model License Map, Read From the Actual Texts
August 8, 2026

What "Open Weights" Lets You Do: The 2026 Model License Map, Read From the Actual Texts

Open Models
The Attention Rebuild: How 2026's Open Models Made 1M Context Fit on Machines You Own
August 1, 2026

The Attention Rebuild: How 2026's Open Models Made 1M Context Fit on Machines You Own

Software & Tools
Speculative Decoding, Explained: The Free Speed Toggle Your Local LLM Is Probably Not Using
July 25, 2026

Speculative Decoding, Explained: The Free Speed Toggle Your Local LLM Is Probably Not Using

Software & Tools
Bandwidth, Not TFLOPS: What Sets Your Local LLM Speed (and Why the Newest Card Isn't Always Fastest)
June 30, 2026

Bandwidth, Not TFLOPS: What Sets Your Local LLM Speed (and Why the Newest Card Isn't Always Fastest)

GPUs for Local LLM
GPT-5.6 is here, and you can't run it. Here's what you can run instead.
June 29, 2026

GPT-5.6 is here, and you can't run it. Here's what you can run instead.

Open Models
Serving a Local LLM as an API: From Ollama's Endpoint to vLLM Throughput (and When to Rent Instead)
June 21, 2026

Serving a Local LLM as an API: From Ollama's Endpoint to vLLM Throughput (and When to Rent Instead)

Software & Tools
RAG on a Local LLM, Explained: Give Your Model Your Documents Without Drowning in Context
June 19, 2026

RAG on a Local LLM, Explained: Give Your Model Your Documents Without Drowning in Context

Software & Tools
Prompt Processing vs Generation: Why Your Box Is Fast at One and Slow at the Other
June 17, 2026

Prompt Processing vs Generation: Why Your Box Is Fast at One and Slow at the Other

Software & Tools
June 15, 2026

The KV Cache, Explained: Why Long Context Eats Your VRAM (and How to Fit More)

Software & Tools
How Much VRAM Do You Actually Need to Run a 70B Model Locally?
June 13, 2026

How Much VRAM Do You Actually Need to Run a 70B Model Locally?

Software & Tools
Ollama vs LM Studio vs llama.cpp: Which Local LLM Runtime Should You Actually Use?
June 12, 2026

Ollama vs LM Studio vs llama.cpp: Which Local LLM Runtime Should You Actually Use?

Software & Tools
Mixture-of-Experts (MoE), Explained: Why “Active Parameters” Decide What Runs on Your Machine
June 11, 2026

Mixture-of-Experts (MoE), Explained: Why “Active Parameters” Decide What Runs on Your Machine

Software & Tools
Your Local AI Model Folder Is a Mess: Taming a Multi-Terabyte Model Hoard on Apple Silicon
June 11, 2026

Your Local AI Model Folder Is a Mess: Taming a Multi-Terabyte Model Hoard on Apple Silicon

Unified-Memory AI
GGUF vs GPTQ vs AWQ: The Plain-English Guide to LLM Quantization (and Which One to Pick)
June 6, 2026

GGUF vs GPTQ vs AWQ: The Plain-English Guide to LLM Quantization (and Which One to Pick)

Software & Tools

Get the Vetted Consumer newsletter

Reviews, buying advice, and field notes. Delivered monthly.

Almost there, check your inbox and click the confirmation link. ✓

Something went wrong, please try again, or email [email protected].