Agreed. Here's a start for posterity. Using their prompt, "Write a 300-word explanation of how attention works in transformer models, aimed at a junior developer. Use one concrete analogy.":
Note that my mini is the bigger Pro model, so it will be faster than the mini quoted in the article.
Mac mini M4 Pro (cores: 10P/4E/20G) 64GB
Tahoe 26.6
LM Studio 0.4.20+1
Qwen3.6-35B-A3B-MLX-4bit: 78.81 tok/s, TTFT 0.93s
(but note Qwen 35B is really chattery and outputs 3,451 words of thinking for 65s first)
Qwen3.6-27B-MLX-4bit: 14.65 tok/s, TTFT 0.76s
(Qwen 27B output 3,027 words of thinking for 361s first, spinning up the fans)
gemma-4-26B-A4B-it-QAT-MLX-4bit: 64.93 tok/s, TTFT 0.44s
(Gemma 26B is much more on-task, thinking with 591 words for 15.62s first)
gemma-4-31B-it-QAT-GGUF Q4_0: 12.25 tok/s, TTFT 1.65s
(Gemma 31B thought with 404 words for 51.55s)
Fair point. The first run covers what our customers deploy most via Ollama today, which skews to the Llama / Qwen 2.5 / Mistral / DeepSeek families. Qwen 3.6 and Gemma 4 are top of the queue for the next run same method, same raw JSON. Which quants would be most useful to you: Q4_K_M only, or Q8 as well?
Why do no benchmarks show qwen3.6?
As far as i can tell the state of the art small open models is qwen3.6 and gemma 4, yet they rarely appear in benchmarks - even ones made recently
Agreed. Here's a start for posterity. Using their prompt, "Write a 300-word explanation of how attention works in transformer models, aimed at a junior developer. Use one concrete analogy.":
Note that my mini is the bigger Pro model, so it will be faster than the mini quoted in the article.
Mac mini M4 Pro (cores: 10P/4E/20G) 64GB Tahoe 26.6 LM Studio 0.4.20+1
Qwen3.6-35B-A3B-MLX-4bit: 78.81 tok/s, TTFT 0.93s (but note Qwen 35B is really chattery and outputs 3,451 words of thinking for 65s first)
Qwen3.6-27B-MLX-4bit: 14.65 tok/s, TTFT 0.76s (Qwen 27B output 3,027 words of thinking for 361s first, spinning up the fans)
gemma-4-26B-A4B-it-QAT-MLX-4bit: 64.93 tok/s, TTFT 0.44s (Gemma 26B is much more on-task, thinking with 591 words for 15.62s first)
gemma-4-31B-it-QAT-GGUF Q4_0: 12.25 tok/s, TTFT 1.65s (Gemma 31B thought with 404 words for 51.55s)
Fair point. The first run covers what our customers deploy most via Ollama today, which skews to the Llama / Qwen 2.5 / Mistral / DeepSeek families. Qwen 3.6 and Gemma 4 are top of the queue for the next run same method, same raw JSON. Which quants would be most useful to you: Q4_K_M only, or Q8 as well?