Category
Software & Tools
Honest guides to the software that runs local LLMs — inference engines, GUIs, quantization, serving, and RAG tooling. Mostly free, open-source tools, covered because they matter to the niche, not for affiliate links.
June 30, 2026
Bandwidth, Not TFLOPS: What Sets Your Local LLM Speed (and Why the Newest Card Isn't Always Fastest)
GPUs for Local LLM
June 29, 2026
GPT-5.6 is here, and you can't run it. Here's what you can run instead.
Open Models
June 21, 2026
Serving a Local LLM as an API: From Ollama's Endpoint to vLLM Throughput (and When to Rent Instead)
Software & Tools
June 19, 2026
RAG on a Local LLM, Explained: Give Your Model Your Documents Without Drowning in Context
Software & Tools
June 17, 2026
Prompt Processing vs Generation: Why Your Box Is Fast at One and Slow at the Other
Software & Tools
June 15, 2026
The KV Cache, Explained: Why Long Context Eats Your VRAM (and How to Fit More)
Software & Tools
June 13, 2026
How Much VRAM Do You Actually Need to Run a 70B Model Locally?
Software & Tools
June 12, 2026
Ollama vs LM Studio vs llama.cpp: Which Local LLM Runtime Should You Actually Use?
Software & Tools
June 11, 2026
Mixture-of-Experts (MoE), Explained: Why “Active Parameters” Decide What Runs on Your Machine
Software & Tools
June 11, 2026
Your Local AI Model Folder Is a Mess: Taming a Multi-Terabyte Model Hoard on Apple Silicon
Unified-Memory AI
June 6, 2026
GGUF vs GPTQ vs AWQ: The Plain-English Guide to LLM Quantization (and Which One to Pick)
Software & Tools