
Qwen3.8-27B for the RAM Poor Mac user:
For those of you that want a functional 24GB Mac laptop while having this overthinking creature boosting your ideas.
With and without MTP (for the desperate)

For those of you that want a functional 24GB Mac laptop while having this overthinking creature boosting your ideas.
With and without MTP (for the desperate)
thanks,
Amazing news!!!
Sony a74
12-24 F4
Seafrogs Housing
Properly done this time.
Models as the GIFs are displayed:
Qwen3.6-27B - 4bit
Qwen3.6-MoE - 6bit
Ornith-35B - 6bit
Gemma-4-26B - 6bit
Qwen3.6-MoE - 4bit
HuiHui-Qwen3.6-MoE - 6bit
Agents-A1 - 6bit
Inference parameters:
Qwen3.6 & HuiHui Abliterated:
temperature 0.6 - top_p 0.95 - top_k 20 - min_p 0.01 - repeat_penalty 1.05
Ornith-1.0-35B:
temperature 1.0 - top_p 1.0 - top_k 40 - min_p 0.01 - repeat_penalty 1.05
Gemma-4:
temperature 1.0 - top_p 1.0 - top_k 64 - min_p 0.01 - repeat_penalty 1.1
Agents-A1:
temperature 0.85 - top_p 0.95 - top_k 20 - min_p 0.01 - repeat_penalty 1.05
Prompt: "Create a beautiful, relaxing flight simulator in a single html file with mountains, clouds, and endless procedural terrain"
Harnes: Pi
Served by: oMLX
Method: single prompt. If the html file doesn't work everything was deleted, Pi session was restarted, and model had to start from scratch again. maximum of 3 tries.
Models Quants used:
https://huggingface.co/collections/leonsarmiento/local-sota-for-48gb-macs
Kind of unexpected.
Happy for Gemma-4/Google, big win for us, LocalLLMers.
Yet Qwen3.6 still does better in Hermes than Gemma-4 somehow.
We need Gemma-4.1 fine-tuned on Agentic-tasks. That would be killer.
“token efficiency needs to drop to as much as 20% over the next 12 months, and 90% by the following year”
That’s it. Need the oMLX to unload and load different models when agent (Hermes) ask for it. It works better with Pi (first time different model is requested returns an error, on second try it unloads the previous and loads the new) but it is stuck with Hermes.
LM studio does this gracefully.