Category
Software & Tools
Honest guides to the software that runs local LLMs — inference engines, GUIs, quantization, serving, and RAG tooling. Mostly free, open-source tools, covered because they matter to the niche, not for affiliate links.
September 4, 2026
Local Embedding Models Explained: The Other Model Your RAG Setup Needs
Software & Tools
September 3, 2026
LLM Sampling Settings Explained: Temperature, Top-P, Top-K, and Min-P
Software & Tools
September 2, 2026
How Much RAM Do You Need to Run a Local LLM in 2026?
Software & Tools
August 27, 2026
The M5 Mac and Local LLMs: Apple's New Matmul Hardware Finally Attacks Prompt Processing
Unified-Memory AI
August 24, 2026
The Best Open LLM You Can Actually Run Right Now, by VRAM Tier (August 2026)
Open Models
August 8, 2026
What "Open Weights" Lets You Do: The 2026 Model License Map, Read From the Actual Texts
Open Models
August 1, 2026
The Attention Rebuild: How 2026's Open Models Made 1M Context Fit on Machines You Own
Software & Tools
July 25, 2026
Speculative Decoding, Explained: The Free Speed Toggle Your Local LLM Is Probably Not Using
Software & Tools
June 30, 2026
Bandwidth, Not TFLOPS: What Sets Your Local LLM Speed (and Why the Newest Card Isn't Always Fastest)
GPUs for Local LLM
June 29, 2026
GPT-5.6 is here, and you can't run it. Here's what you can run instead.
Open Models
June 21, 2026
Serving a Local LLM as an API: From Ollama's Endpoint to vLLM Throughput (and When to Rent Instead)
Software & Tools
June 19, 2026
RAG on a Local LLM, Explained: Give Your Model Your Documents Without Drowning in Context
Software & Tools
June 17, 2026
Prompt Processing vs Generation: Why Your Box Is Fast at One and Slow at the Other
Software & Tools
June 15, 2026
The KV Cache, Explained: Why Long Context Eats Your VRAM (and How to Fit More)
Software & Tools
June 13, 2026
How Much VRAM Do You Actually Need to Run a 70B Model Locally?
Software & Tools
June 12, 2026
Ollama vs LM Studio vs llama.cpp: Which Local LLM Runtime Should You Actually Use?
Software & Tools
June 11, 2026
Mixture-of-Experts (MoE), Explained: Why “Active Parameters” Decide What Runs on Your Machine
Software & Tools
June 11, 2026
Your Local AI Model Folder Is a Mess: Taming a Multi-Terabyte Model Hoard on Apple Silicon
Unified-Memory AI
June 6, 2026
GGUF vs GPTQ vs AWQ: The Plain-English Guide to LLM Quantization (and Which One to Pick)
Software & Tools