Image 1 — Something wrong with my Qwen3.8 27B local speed on MacBook 128G
Image 2 — Something wrong with my Qwen3.8 27B local speed on MacBook 128G
Image 3 — Something wrong with my Qwen3.8 27B local speed on MacBook 128G
▲ 9 r/oMLX

Something wrong with my Qwen3.8 27B local speed on MacBook 128G

I'm using omlx launch opencode, with the Qwen3.8-27B-oQ8e-fp16-mtp checkpoint, on a MacBook M5 Max (40c) 128GB. oMLX version is 0.6.3rc1. All configurations are default; only context length is 262K.

Sorry, I may not be so familiar with using oMLX. Why is it so slow? What should I do to optimize it?

Also, how should I make use of the z-lab/Qwen3.8-27B-DFlash2 checkpoint to further speed it up? It does not load standalone in Models; how should I make it load with the Qwen 3.8 checkpoint?

u/Altruistic-Dust-2565 — 21 hours ago
▲ 11 r/oMLX

Does the new ds 0731 oQ2e mtp fit in 128G Mac?

The new 2bit deepseek is just uploaded https://huggingface.co/Jundot/DeepSeek-V4-Flash-0731-oQ2e-mtp.

I am used to download model from LM Studio (so that both LM Studio and oMLX can recognize. Presumably they are from the same shared huggingface source. Would downloading from LM Studio affect how oMLX recognize this model for MTP? Or must I download from oMLX?

Also, more importantly it is showing "Partial GPU offload possible" in LM Studio, would it be the same for oMLX? It is only 106G, but my 128G Macbook is shown to have only 105G VRAM somehow.

Also what about context length? I aim for 262K minimal and full 1M if possible. Deepseek has low kv cache cost if I'm understanding it correctly?

In a nutshell, how should I run oQ2e safely on my 128G Mac? Should I drop MTP to free more VRAM for full GPU offload? Or wait for a smaller oQ? Or use this version with a low context? Or does everything already fit?

u/Altruistic-Dust-2565 — 19 days ago
▲ 7 r/accelerate+1 crossposts

$200 ChatGPT just saved me $2000 on MacBook purchase

I have a 5-year-old M1 Pro and had been waiting for the M5 Ultra Mac Studio upgrade for a year. After WWDC I was devastated and started considering a MacBook Pro instead.

After the price increase I rushed to look for deals, but found nothing. Just as I was about to completely give up, with nothing to lose, I asked ChatGPT Pro (Extended) to search all available online stores for a 16-inch MacBook Pro M5 Max (18+40), 128GB+2TB, nano-texture display, and find the lowest price.

After a while it gave me a link that seemed to have the pre-price-increase price. At first I thought it had hallucinated it, and even explained the recent price increase to it, because when I opened the link it still showed the new higher price.
But it turned out the seller (one of Apple’s official retailers on JD.com, basically the Chinese Amazon) had somehow FORGOTTEN to update the price for that exact configuration… and it just happened to be the one I wanted.

I immediately placed the order, and within 8 hours I had my brand-new MacBook for about ¥43,000 (~US$6,329), while the official Apple China Store charges ¥57,124 (~US$8,408), and the cheapest unofficial sellers are around ¥49,000 (~US$7212) because of import taxes and such.

I’m still amazed that ChatGPT managed to find the one glitched configuration.

For the record, switching from 128GB to 64GB actually increased the price. Removing the nano-texture display or even applying the education discount also made it more expensive. Only that exact combination had the glitched price.

I’ll never complain that AI is getting too expensive again… but I’ll definitely keep complaining that Apple products and RAMs are. XD

Hope all of you can find your desired Macs!

reddit.com
u/Altruistic-Dust-2565 — 2 months ago

What's the best frontend dev setup after Fable

I have a small personal project that's halfway done. Python backend + react frontend (purely vibe coded frontend via Codex GPT-5.5).

Would Fable + Claude Code be able to maintain and remarkbly improve it autonomously over the GPT-5.5 version? If so, do I use it for planning only or a full frontend upgrade refactor?

What's the best set of tools? Playwright? Kernel? or screenshot and computer use?

Do I use loop or workflow or ultrawork or something?

Should I let it follow current implementation, or how do I prompt it for an upgrade or a complete redesign?

I have been using Codex for too long, and as a pure Python engineer it's like only my second frontend project ever so I have absolutely no idea.

FYI I have Codex $200 and Cursor $20, consider adding another Claude Code $20 on top of them and still keeping Codex my main worker I guess?

reddit.com
u/Altruistic-Dust-2565 — 2 months ago

So in no way would I quit Annual Pro+ now

Started my annual pro+ last-minute in March.

I guess I'll have to stick to it until GPT-6 comes out.

u/Altruistic-Dust-2565 — 3 months ago

Why there isn't any top LLM providers investing on diffusion LLM?

A year ago, I would’ve said Diffusion LLMs were an interesting idea but still far from practical. They’re still pretty rough, but Mercury 2 now makes it seem like they might finally be getting close to usable.

That said, aside from Meta, Ant, and Inception/Mercury, it doesn’t seem like many labs are seriously investing in them — especially the major ones like OpenAI, Anthropic, Google, xAI, or even architecture-focused teams like DeepSeek and Kimi.

I’m not very familiar with DLLMs, so I’m curious: why is that? Are there still fundamental issues with the paradigm that make them unlikely to become even second-tier models? Or is current hardware stack a bottleneck for DLLMs training/inference? Or are other labs just working on it quietly and not there yet?

reddit.com
u/Altruistic-Dust-2565 — 3 months ago

"Early May" is ending, where is the preview?

BTW, what happens if I don’t ask for a refund by May 20? Can I still request it in June? I would really like to experience the new multipliers first and then decide.

I’m still undecided about switching from Copilot to Codex, because a few things are unclear:

  1. The Copilot GPT-5.5 model multiplier still hasn’t been disclosed. (https://www.reddit.com/r/GithubCopilot/comments/1t53aqr/copilot_gpt55_multiplier_is_now_listed_as_75x_tbd/?utm_source=share&utm_medium=web3x&utm_name=web3xcss&utm_term=1&utm_content=share_button)

  2. There’s no usage preview — I need token usage visibility to compare it properly with Codex. (tried AI Engineering Fluency plugin, which counts the tokens but seems to miss a lot).

  3. The refund policy is vague. I'm not sure what percentage will be refunded. Has anyone received theirs yet? How was it calculated? remaining days / 365? How long does it take? Can I file the request on May 19th?

u/Altruistic-Dust-2565 — 3 months ago

What the hell does TBD even mean here?

Copilot, are you seriously saying you still haven’t decided how much GPT-5.5 — which has been out for two weeks now — is going to cost? Because this basically reads like:

> “We’ve already decided we’re charging you more, we just haven’t figured out exactly how much more we can squeeze out of you yet.”

At least for now, I guess we can entertain the fantasy that maybe some new specialized chips will roll out (like when Cerebras powered Codex-Spark), and GPT-5.5 pricing could actually come down due to newer deployments. Or maybe Microsoft and Sam Altman are in the middle of some other negotiations right now?

u/Altruistic-Dust-2565 — 4 months ago