Qwen3.8-27B for the RAM Poor Mac user:
▲ 15 r/oMLX+1 crossposts

Qwen3.8-27B for the RAM Poor Mac user:

For those of you that want a functional 24GB Mac laptop while having this overthinking creature boosting your ideas.

With and without MTP (for the desperate)

huggingface.co
u/JLeonsarmiento — 1 day ago
🔥 Hot ▲ 15.3k r/damnthatssobering+7 crossposts

Selected as one of the finalist of Ocean Photographer of the Year 2026!!

Amazing news!!!

Sony a74
12-24 F4
Seafrogs Housing

u/JLeonsarmiento — 3 days ago

FlightSimulatorBench: Small MoE edition

Properly done this time.

Models as the GIFs are displayed:

  1. Qwen3.6-27B - 4bit

  2. Qwen3.6-MoE - 6bit

  3. Ornith-35B - 6bit

  4. Gemma-4-26B - 6bit

  5. Qwen3.6-MoE - 4bit

  6. HuiHui-Qwen3.6-MoE - 6bit

  7. Agents-A1 - 6bit

Inference parameters:

Qwen3.6 & HuiHui Abliterated:

temperature 0.6 - top_p 0.95 - top_k 20 - min_p 0.01 - repeat_penalty 1.05

Ornith-1.0-35B:

temperature 1.0 - top_p 1.0 - top_k 40 - min_p 0.01 - repeat_penalty 1.05

Gemma-4:

temperature 1.0 - top_p 1.0 - top_k 64 - min_p 0.01 - repeat_penalty 1.1

Agents-A1:

temperature 0.85 - top_p 0.95 - top_k 20 - min_p 0.01 - repeat_penalty 1.05

Prompt: "Create a beautiful, relaxing flight simulator in a single html file with mountains, clouds, and endless procedural terrain"

Harnes: Pi

Served by: oMLX

Method: single prompt. If the html file doesn't work everything was deleted, Pi session was restarted, and model had to start from scratch again. maximum of 3 tries.

Models Quants used:

https://huggingface.co/collections/leonsarmiento/local-sota-for-48gb-macs

u/JLeonsarmiento — 29 days ago

Updated Gemma-4 chat template witchcraft: Gemma-4-26B-a4B shows dominance over Qwen3.6-MoE and Qwen3.5-MoE fine tunes (Instruct mode and Reasoning efficiency)

Kind of unexpected.

Happy for Gemma-4/Google, big win for us, LocalLLMers.

Yet Qwen3.6 still does better in Hermes than Gemma-4 somehow.

We need Gemma-4.1 fine-tuned on Agentic-tasks. That would be killer.

u/JLeonsarmiento — 29 days ago

CEO: “token efficiency needs to drop 90%” Dude… just write “\no_think” before you ‘summarize this email’ prompts

“token efficiency needs to drop to as much as 20% over the next 12 months, and 90% by the following year”

cnbc.com
u/JLeonsarmiento — 1 month ago
▲ 3 r/oMLX

How to “kept only last loaded model” option, a la LM studio?

That’s it. Need the oMLX to unload and load different models when agent (Hermes) ask for it. It works better with Pi (first time different model is requested returns an error, on second try it unloads the previous and loads the new) but it is stuck with Hermes.

LM studio does this gracefully.

reddit.com
u/JLeonsarmiento — 2 months ago
▲ 2.9k r/InesperadoCu+4 crossposts

The CIA rectal tool kit for emergency, full of useful tools for escape or defense. (1960s-1980s)

u/Aloha-Eh — 2 months ago