
Any luck with dflash2?
Has anyone had luck with using dflash2 for Qwen3.8 27B?
I tried on my M5 Max and did not notice much of an improvement over lightning mtp

Has anyone had luck with using dflash2 for Qwen3.8 27B?
I tried on my M5 Max and did not notice much of an improvement over lightning mtp
I’m on an M5 Max 128gb and get the below error
Model: Vontra/Ling-3.0-flash-MLX-4bit
q_state, k_state, v_state, recurrent_state = cache
ValueError: not enough values to unpack (expected 4, got 0)
Anyone know what the issue might be?
The Ling-3.0-flash MoE is now open-weighted at 124B A5B params. I know the original announcements were before the Kimi K3, DeepSeek-V4-Flash and Qwen3.8 hype, but this model might still have a good niche for itself due to its sizing.
Discussion on the benchmarks are here: https://www.reddit.com/r/LocalLLaMA/comments/1v4mltt/benchmarks_antling30flash_a_hybridreasoning_moe/ from almost 2 weeks ago.
Recently started using the pro version of Chapper with local models. Using qwen3.5 0.8B, 1.7B, and 4B all MLX 4 bit.
Lots of crashes are happening. Whenever I try to enable repeat penalty no matter the value it crashes the app during chat, whenever I try to enable really any of the custom parameters it nets to a crashed app.
Also, even after getting through a couple rounds of prompts with one of the above models it almost always ends up crashing the app. Tool calls are also extremely glitchy.
At this point, I find the app completely unusable beyond an experimental first prompt.
iPhone 17 Pro Max is the device I’m using.