▲ 1 r/oMLX

Can't Quantize Muse-Glimmer-30B

I was able to quantize the following models on my M3Max 64GB with 58GB allocated to max GPU limit:

  • Qwen3.8-27B-oQ8e-mtp (Original: 55.59 GB)
  • Qwen3.6-35B-A3B-oQ8e-mtp (Original: 71.93 GB)
  • gemma-4-31B-it-oQ8e (Original 62.58 GB)

Muse-Glimmer-30B (Original: 59.58 GB) is smaller than Qwen3.6-35B and gemma-4-31B. However, I get out of memory error when I try to quantize Muse-Glimmer-30B.

  • OQ LEVEL: oQ8e
  • Enhanced quantization: On
  • Text Only: Off

I'm including the error below. I'd appreciate any suggestion.

Thanks!

2026-08-16 00:36:56,868 - omlx.admin.oq_manager - ERROR - oQ quantization failed: Muse-Glimmer-30B -> oQ8: auto-proxy sensitivity failed ([METAL] Command buffer execution failed: Insufficient Memory (00000008:kIOGPUCommandBufferCallbackErrorOutOfMemory).). Pass sensitivity_model_path with a pre-quantized version of this model, or run on a machine with enough RAM for full-fp16 sensitivity measurement.
Traceback (most recent call last):
  File "/Volumes/Samples/ai/omlx/omlx/oq.py", line 5716, in quantize_oq_streaming
    sensitivity_map = _measure_sensitivity_from_quantized_model(
        str(_proxy_dir),
    ...<4 lines>...
        trust_remote_code=trust_remote_code,
    )
  File "/Volumes/Samples/ai/omlx/omlx/oq.py", line 8340, in _measure_sensitivity_from_quantized_model
    mx.synchronize()
    ~~~~~~~~~~~~~~^^
RuntimeError: [METAL] Command buffer execution failed: Insufficient Memory (00000008:kIOGPUCommandBufferCallbackErrorOutOfMemory).

The above exception was the direct cause of the following exception:

Traceback (most recent call last):
  File "/Volumes/Samples/ai/omlx/omlx/admin/oq_manager.py", line 611, in _run_quantization
    await asyncio.to_thread(
    ...<19 lines>...
    )
  File "/Users/cgk/.pyenv/versions/3.13.14/lib/python3.13/asyncio/threads.py", line 26, in to_thread
    return await loop.run_in_executor(None, func_call)
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/Users/cgk/.pyenv/versions/3.13.14/lib/python3.13/concurrent/futures/thread.py", line 59, in run
    result = self.fn(*self.args, **self.kwargs)
  File "/Volumes/Samples/ai/omlx/omlx/oq.py", line 5725, in quantize_oq_streaming
    raise RuntimeError(
    ...<4 lines>...
    ) from e
RuntimeError: oQ8: auto-proxy sensitivity failed ([METAL] Command buffer execution failed: Insufficient Memory (00000008:kIOGPUCommandBufferCallbackErrorOutOfMemory).). Pass sensitivity_model_path with a pre-quantized version of this model, or run on a machine with enough RAM for full-fp16 sensitivity measurement.
reddit.com
u/chibop1 — 4 days ago

Less Than a Month: Kimi K3, Qwen3.8, DeepSeek-V4-Pro-0813, GLM-5.3

What's happening in China?

  • Kimi K3-2.8T
  • Qwen3.8-2.4T
  • DeepSeek-V4-Pro-0813-1.6T
  • GLM-5.3-743B

They’re all less than a month old!

reddit.com
u/chibop1 — 6 days ago
▲ 32 r/comfyui

Which Turbo Lora for Minimax-H3 on ComfyUI?

Sorry, newbie here.

I see a lot of posts about Turbo Lora.

Which one is the best to use right now, and where do I get them?

Thanks!

reddit.com
u/chibop1 — 12 days ago
▲ 53 r/AudioAI+1 crossposts

Parlor v2: best-effort fully local GPT-Live clone on an M3 Pro

GPT-Live is so good that I use it almost every day. I've been wanting to replicate it since it was released.

My first attempt was to fine-tune Gemma 4 12B to behave like a full-duplex model. Something like grafting a decision tick + speech head to the model. It failed after multiple trials. For now, I think a classic cascade system is still better. We just need to wait until a benevolent frontier AI company releases a full-duplex model that's on par with GPT-Live.

Repo: https://github.com/fikrikarim/parlor/

u/chibop1 — 18 days ago

People reported @hypabolic/pi-hypa 4 times, but just shot down?

Am I just being overly cautious?

@hypabolic/pi-hypa is #1 on the chart, with 203K monthly downloads and 191.1K weekly downloads. In comparison, pi-web-access, which is #2 on the chart, has 113.9K monthly downloads and 32K weekly downloads.

The GitHub repo for pi-hypa has only 34 stars, 1 fork, and 0 watchers.

There are four package reports, including mine, on earendil-works/pi, but they all seem to have been closed without a concrete answer, other than that they do not have the bandwidth to check each package.

Maybe it is worth investigating the #1 package on their chart?

Or maybe it really is that useful and popular, and there is nothing wrong with the package.

? I installed, tried, and uninstalled it. Am I paranoid?

u/chibop1 — 2 months ago

Pi Can't Access pi-hypa on Windows?

It looks like @hypabolic/pi-hypa shot up to #1 extension with most monthly download, so I wanted to try to replace context-mode which adds like almost 10k tokens!

I installed with pi install npm:@hypabolic/pi-hypa from Windows command line.

I'm on Pi 0.79.10.

It seems like it's installed.

/hypa
Hypa Pi extension
Mode: additive
Binary: hypa
Resolved binary: C:\Users\<username>\.pi\agent\npm\node_modules\@hypabolic\hypa\bin.js
Rewrite timeout: 5000ms
Ask fallback (non-UI): deny
MCP proxy discovery: disabled
MCP proxy timeout: 10000ms
Pi MCP config for dedup: default
Active Hypa tools: hypa_shell, hypa_read, hypa_grep, hypa_find, hypa_ls
Last rewrite: none

Then I asked an agent (gpt-5.5) to run Hypa commands, but the agent couldn't run any of them.

User: Can you run hypa tools?
Assistant: Yes. I can run Hypa tools here, including:

- hypa_shell
- hypa_read
- hypa_grep
- hypa_find
- hypa_ls

Tell me what you want to inspect or run.

User: Can you test running some hypa tools?

Tool: hypa_ls
spawn EFTYPE

Tool: hypa_find
spawn EFTYPE

Troubleshooting executable format

I encountered an issue with an unsupported executable format, denoted by EFTYPE. Perhaps I should run a report and see if trying a different approach might help? It might also be worth testing with
hypa_shell/read to gather more insight. I need to figure out the next steps to resolve this problem efficiently. There’s a bit of trial and error here, but I'm determined to find a solution!

Tool: hypa_shell
spawn EFTYPE

Tool: hypa_grep
spawn EFTYPE

I tested several Hypa tools:

- hypa_ls
- hypa_find
- hypa_shell
- hypa_grep

All returned the same error:

```text
  spawn EFTYPE
```

So the Hypa tools are available in the environment, but currently failing to execute—likely an executable/platform/type issue with the Hypa binary.

Does anyone got it to work on Windows command line?

Thanks!

reddit.com
u/chibop1 — 2 months ago

Your Favorite Workflow to Convert PDF with Complex Structure to Markdown?

I've tried markitdown, Docling, and Mineru.

Are there better tools I should try?

I need to process tables, floating box, etc.

Thanks!

reddit.com
u/chibop1 — 2 months ago

Background Tool Execution?

Can Pi run a long-running tool in the background, or is there an extension for it?

Claude Code can launch a task in the background, allowing to keep chatting while it runs. When the tool finishes, it brings back the result into the conversation and continues the discussion seamlessly.

I don't need to fork or create new session or anything.

In Pi, I have to wait until the tool call finishes, and sometimes it could take a long time for tasks like heavy analysis tool.

Thanks!

reddit.com
u/chibop1 — 2 months ago
▲ 42 r/openwrt

Is Flint 2 Still Best OpenWRT Router in 2026?

I'd like to upgrade my home router with WiFi. Is Flint 2 Still the Best OpenWRT Router?

What about 6ghz?

I live in a city with many apartment buildings and units, so there are countless Wi-Fi networks nearby.

Thanks!

reddit.com
u/chibop1 — 2 months ago

favorite Agentic Coding Harness

So far, I’ve tried Codex CLI, Claude Code, Gemini CLI, OpenCode, and recently, Pi with local models.

Pi is the leanest of them all, with just four tools: read, write, edit, and bash. Its system prompt is only under 2K tokens, and it's perfect for local models.

I've been trying out Qwen 27B-MXFP8 with it, and it's much better than I expected!

It doesn't have fancy bells and whistles like multi agents, but the only thing I’m missing is searching the web for documentation. I’m sure you can get it through an extension, but you probably won’t get the same robust search features you get from commercial platforms anyways.

This might be my new favorite! What’s yours?

reddit.com
u/chibop1 — 3 months ago
▲ 21 r/AudioAI+1 crossposts

Convert With MPT Support?

Hi All,

I'm trying to understand the process of creating GGUF with MTP support.

Does the original Qwen/Qwen3.6-27B support MTP?

If not, how do you revise the original model to support MTP?

Also, is there a special flag I need to use to convert that into GGUF to retain the MTP capability?

Thanks!

u/chibop1 — 6 days ago