Editor checklist: every speed below is someone's measurement, attributed and linked, except the ceilings, which are bandwidth arithmetic and labelled as such. The only first-party figure is our own Mac mini's llama.cpp run. We have not run MLX on it yet. Delete this box before publishing.
If you run models on a Mac, you are choosing between two engines whether you know it or not. llama.cpp is the cross-platform C++ runtime behind GGUF files, LM Studio's default engine and, until this year, Ollama on every platform. MLX is Apple's own array framework, built for unified memory, with its own model format and a growing library of converted models. The usual advice is "MLX is faster on a Mac". The measurements say that is true for generation at short context, with a margin that shrinks or reverses on long prompts. This piece puts the numbers side by side and explains where each engine's lead comes from.
The short answer
| Your workload | Faster engine, per current measurements |
|---|---|
| Chat and short prompts on any Apple Silicon Mac | MLX, by roughly 20–40% in generation |
| Agentic coding or documents at 30K+ tokens of context | llama.cpp with flash attention; MLX has been measured at about half its generation speed at very long context |
| M5-generation Macs | MLX; Apple's published neural-accelerator gains were measured with it |
| One model file shared between a Mac and a PC | llama.cpp (GGUF runs everywhere) |
| Best quality per byte at 4-bit | GGUF Q4_K_M, whose mixed bit allocation is built for exactly that |
Why MLX usually wins generation
On any Mac, generating a token means reading the model's weights from memory once, so speed is capped by memory bandwidth divided by the bytes read (the bandwidth-not-TFLOPS rule). Two things separate the engines at this step.
The first is file size. MLX's default 4-bit format is uniform: in the MLX source, "affine" quantization defaults to groups of 64 weights at 4 bits, each group carrying a 16-bit scale and a 16-bit bias. That works out to 4.5 bits per weight. GGUF's Q4_K_M is mixed: the k-quants pull request that introduced it uses 6-bit precision for half of the attention value and feed-forward output tensors and 4-bit for the rest. Our own Qwen3 8B Q4_K_M file is 5.02GB for 8.19 billion parameters, or 4.9 bits per weight. So an MLX 4-bit file of the same model is about 8% smaller, and 8% fewer bytes per token is up to 8% more speed before any kernel is written.
The second is how much of the bandwidth each engine turns into tokens, and here MLX has the edge. On a Mac Studio M3 Ultra, TerminalBytes measured Qwen3.5 9B on 20 September 2026:
| Runtime | Weights | Generation, tok/s | Prompt processing |
|---|---|---|---|
| llama.cpp (llama-bench, build 10809) | GGUF Q4_K_M | 77.8 | 1,010 tok/s |
| Ollama 0.32.13, GGUF | Q4_K_M | 77.7 | 344 tok/s |
| Ollama 0.32.13, MLX | NVFP4 | 90.5 | 26–37 tok/s |
| mlx-lm 0.31.3 | MLX 4-bit affine | 107.5 | not reported |
That is a 38% generation lead for mlx-lm over llama.cpp on the same model. The 819GB/s M3 Ultra puts rough ceilings near 150 tokens per second for the Q4_K_M file and 160 for the MLX file, so llama.cpp is converting about half its bandwidth into tokens and MLX about two thirds. File size explains a fraction of the gap; the rest is MLX's kernels. The author flags the catch: the MLX and GGUF weights are different quantizations, so the comparison measures a package (format plus engine), which is the choice a user makes anyway.
The academic comparison agrees on direction. Rajesh et al. (2025, arXiv:2511.05502) tested five runtimes on an M2 Ultra with 192GB across the Qwen 2.5 family and prompts up to 100,000 tokens, and found that "MLX achieves the highest sustained generation throughput", while llama.cpp "is highly efficient for lightweight single-stream use" and MLC-LLM had lower time to first token at moderate prompt sizes.
Where the answer flips: long context
Generation at depth is a different workload. Every new token reads the whole KV cache as well as the weights (see the KV cache explainer), and how efficiently the attention kernel reads that cache starts to dominate. An M3 Ultra owner filed mlx-lm issue 763 in January 2026 with MiniMax M2.1 at 4-bit, run through LM Studio on both engines:
| Context | MLX generation | llama.cpp (flash attention) generation |
|---|---|---|
| 30K tokens | 25 tok/s | 32 tok/s |
| 146K tokens | 5.95 tok/s | 12.12 tok/s |
Prompt processing was close (82.5 seconds for MLX against 78 for llama.cpp at 30K), but generation at 146K ran at half speed on MLX. The issue was still open when we checked. For agentic coding tools, which carry tens of thousands of tokens of context on every turn, this is the number that matters, and it is the opposite of the short-prompt result. Test your own context length before switching engines for this work.
Prompt processing, and the M5 change
Prefill, reading your prompt before the first token, is compute-bound rather than bandwidth-bound (why the two phases differ). Here the picture depends on the chip generation. On M1 to M4, llama.cpp's Metal kernels hold up well: 1,010 tok/s in the M3 Ultra run above, and 223.8 tok/s on our own Mac mini M4 with Qwen3 8B. The same run measured 19.7 tok/s generation against a ceiling of about 24, so llama.cpp is not leaving much on the table on a base M4 either.
The M5 changed prefill. Apple added matrix-multiply units to every GPU core, and Apple Machine Learning Research reported (19 November 2025) that MLX on an M5 MacBook Pro reached time-to-first-token speedups of 3.33× to 4.06× over an M4 on Qwen 8B and 14B at 4-bit and Qwen 30B MoE, while generation improved only 1.19× to 1.27×, in line with the bandwidth rising from 120GB/s to 153GB/s. The article is explicit about the split: first-token generation "is compute-bound", later tokens are "bounded by memory bandwidth". Those gains were measured with MLX, which is the engine Apple builds for first. Our M5 prompt-processing piece covers the hardware.
Ollama now sits on both sides
Ollama used to mean llama.cpp. On 30 March 2026, Ollama announced an MLX backend for Apple Silicon in preview. On an M5-series Mac with Qwen3.5-35B-A3B at NVFP4, prefill went from 1,154 to 1,810 tok/s and decode from 58 to 112 tok/s compared with the previous llama.cpp-based version. The preview required more than 32GB of unified memory and supported a single model at launch. Two cautions follow. First, check which backend a given Ollama model tag uses; the M3 Ultra table above shows MLX and GGUF tags of the same model behaving very differently. Second, a wrapper can cost more than the engine choice: in Ollama issue 14861, a user measured raw llama.cpp at about 53 tok/s against Ollama's 32 tok/s on the same Qwen3.5 35B file. Our Ollama vs LM Studio vs llama.cpp guide covers the ergonomics.
Quality: the part benchmarks skip
Speed comparisons at "4-bit" hide a quality difference. Q4_K_M exists because uniform 4-bit loses more than it needs to: in the k-quants pull request, a 7B model at Q4_K_M measured a wikitext perplexity of 5.9601 against 5.9066 at full 16-bit precision, a 0.9% increase, by spending extra bits on the tensors most sensitive to rounding. MLX's default affine 4-bit spreads its bits evenly. We could not find a like-for-like perplexity test of MLX 4-bit against Q4_K_M on the same model that we could verify, so we will not put a number on the gap. If quality matters more than the last 20% of speed, MLX's 6-bit and 8-bit conversions, or GGUF Q5_K_M and above, sidestep the question at a cost in memory. The quantization guide explains the formats, and MLX also supports MXFP4 and NVFP4 modes for models that ship in them.
Which one to use
- You chat, summarise short documents, or run a coding assistant with modest context. Use MLX, through LM Studio's MLX engine, mlx-lm, or Ollama's MLX models. Expect roughly 20–40% more tokens per second than the GGUF of the same model.
- You run agents that keep 30K+ tokens in context. Benchmark both at your real context length. On the evidence available, llama.cpp with flash attention holds generation speed better as the context grows.
- You bought an M5-generation Mac. Start with MLX, because Apple's neural-accelerator gains were published on it.
- You have a base Mac with 16–24GB. The engine matters less than fitting the right model. At 20 tokens per second, a 20% gain is four tokens. Pick whichever format your model ships in first, and check fit in the Quant Picker.
- You share models between a Mac and a PC, or care most about quality per byte. Stay with GGUF.
Limits of this comparison
The head-to-head numbers come from a small set of machines and models, and both engines change monthly. The M3 Ultra comparison uses different quantizations for each engine, the long-context result is one owner's report on one model, and Ollama's MLX figures are the vendor's own. We have measured llama.cpp on our Mac mini but not yet MLX, so we cannot add a matched first-party row. Treat the direction of each result as solid and the percentages as approximate.
Sources and how we researched this
Research: Rajesh et al., Production-Grade Local LLM Inference on Apple Silicon: A Comparative Study of MLX, MLC-LLM, Ollama, llama.cpp, and PyTorch MPS (arXiv:2511.05502, 2025); Apple Machine Learning Research, Exploring LLMs with MLX and the Neural Accelerators in the M5 GPU (2025). Primary project sources: MLX quantization defaults, llama.cpp k-quants pull request, Ollama's MLX announcement. Measurements: TerminalBytes' M3 Ultra comparison, mlx-lm issue 763, Ollama issue 14861, and our own Mac mini M4 llama.cpp benchmark. Ceilings are bandwidth arithmetic, not measurements. Non-commercial site, no affiliate links.
Related: Which Mac to buy for local LLMs · M5 Max vs M5 Ultra