Qwen 3.8 accuracy benchmark
Hi I run qwen 3.8 oq4e version and accuracy benchmark of MathQA is only 41% and LiveCodeBench is 52%. Not sure if it’s normal? They look low to me… recall 3.6 has slightly higher score on both. Maybe quantisation issue?
Hi I run qwen 3.8 oq4e version and accuracy benchmark of MathQA is only 41% and LiveCodeBench is 52%. Not sure if it’s normal? They look low to me… recall 3.6 has slightly higher score on both. Maybe quantisation issue?
Try to run it on my m4 max MacBook, the memory usage goes up and down and at one point the whole process stalled so the number of tokens stayed at around 50k for very long time like 15minutes, and during this period the memory usage (oMLX) goes up and down for several gbs. Not sure I understand how it works? Also I have experience the measurement just aborts by itself during the test…
Hi everyone,
My 2022 Model 3 RWD has developed a strange noise from the right side of the front suspension, mainly when driving over speed bumps.
The car was serviced by Tesla in March, when they replaced the FRONT LOWER COMPLIANCE LINK ASSEMBLY – RIGHT HAND (2188359-10-B).
Could the previous repair be related to this noise? Does anyone know what might be causing it?
Video attached. Thanks!
I'm testing Qwen3.6-35B-A3B-oQ4e-mtp on an M4 Max 36GB using oMLX.
I noticed something interesting:
I expected a much larger increase because TurboQuant compresses the KV cache. Since FP16 KV cache is 16-bit and TurboQuant 3-bit should reduce it significantly, I thought the available context window might almost double.
However, I only see about a 10% increase in maximum context. Is it expected? If so any other way to increase context size?
PS. 27B dense model can't give more than 30k so not good enough for coding hence forced to use MoE mode.
I am using Qwen3.6-35B-A3B-oQ4e-mtp with oMLX on an M4 Max 36GB.
I have two questions:
​
TurboQuant KV: ON
KV quantization: q3
Lightning MTP: OFF
Does running an MTP model with MTP disabled cause any problems?
Many thanks!
Hi all, new to oMLX and would like some help.
I have Apple Sillicon m4 max with 36GB ram, would like to make localLLM for coding tasks therefore need bigger context window (100k preferred). With 27B and 35B models I cannot push for more than 75k, hence I need a way to compress KV. I used to run LM Studio with mlx models but it doesn’t allow me to quantisation KV for Qwen models (not sure why but loading the model failed). Now I switch to oMLX and really like it, and even better I find TurboQuant which can compress the KV with less impact on the model. So my questions are:
Hi I just got a new MacBook Pro with base M4 Max chip 14c CPU and 32c GPU. I have tried to run cinebench with both version 2024 and 2026, but consistently get lower scores. I have stopped all applications and connect the laptop to the 96w MagSafe charger, set high power mode, but still the result is lower than reference. My single core result is 159, multi-core is 1580, GPU is 13500.
I noticed for single core test cinebench or Mac OS runs on multiple cores and it never really stressed the core and speed is always around 2ghz.
For multiple cores the performance cores all all stressed and during 10minutes testing keeps between 3.2ghz and 3.6ghz. It never reaches the peak 4.5Ghz though, not sure if expected.
I understand m4 max can easily score over 175 single core and 1780 multi-core, so my configuration is like 10-20% slower, not sure where to check?
PS: for multi-core the cpu stays at high frequency so no throttling.
Hi all:
Would like to get some help with local LLM for coding tasks. I have a MacBook Pro with M4 Max chip and 36G ram. I have tried Qwen 3.5 27B 4bits MLX with LM studio, it works with token generation speed around 10-15/second. I’d like to use the localLLM for coding tasks, I have a hobby Python + React/TypeScript project with several thousand lines of code. Would like to ask:
Thank you very much in advance!
Just bought it before the price hike, under normal use suddenly black screen, not possible to turn on or charge after that. Sent it back 2nd day to the retailer, and they sent it to Apple authorized repair shop. Latest news I heard: they replaced the logic board and the device is now on its way back.
This is my only Apple device that failed within hours, and the webshop I bought it from can’t replace it as there is no more m4 MacBook Pro left.
Should I be worried that they had to replace the logic board? And am I still eligible to add AppleCare after it’s repaired?
I have been using Emby for over 10 years, in the past 2-3 years I feel there was little changes, no new functionalities or improvements. Anybody know if there is a roadmap?
It’s insane
I’m looking for a 14" MacBook Pro for personal use. Main use cases are Lightroom/photo editing, normal daily use, some coding, and experimenting with local LLMs for coding.
I’m choosing between two Apple preconfigured models:
The price difference is roughly €750.
For storage, I think 1TB is probably enough because I have a NAS, and photos are only kept locally while editing/sorting. My bigger question is whether the M5 Max is worth the extra money, especially for Lightroom and local LLM/coding use.
Since both have the same 36GB RAM, I’m wondering if the M4 Max is the better value, or if the M5 Max is worth it for the newer chip, extra CPU cores, 2TB SSD, and resale value.
I’d also like to run something like Qwen3.6 27B/Qwen2.5-Coder 32B locally for coding assistance. Is that realistic with 36GB unified memory?
Apple also claims big AI performance improvements — in some cases up to around 4x. For real local LLM usage, would the M5 Max actually feel anywhere near 4x faster than the M4 Max, or is that mostly for specific benchmark/Apple Intelligence workloads?
Which one would you buy?
Hi:
I’m seriously considering to buy a couple of FP300 sensors to control lights. I don’t have Aqara ZB 3.0 hub, neither do I have any other ZB dongles. My home automation system Home Assistant and Apple HomeKit. I’m considering to use matter mode but read some conflicting reports on functionalities exposed to Apple HomeKit, so hopefully somebody use matter mode can help me:
Running cronjob as isolated from sub-agent, it worked fine yesterday after upgrade I got the error:
Cannot complete the cron payload from this environment because I do not have workspace shell or filesystem access to
OpenClaw finding:
Fresh manual run:
dfbebf5f-8a62-4882-a507-cfec6fd72c1fmanual:dfbebf5f-8a62-4882-a507-cfec6fd72c1f:1779366702437:1d9ec7cc5-0b6e-4ef1-9068-5e04b4637322gpt-5.4-mini via openai-codexWhat happened in the new session:
toolCount was still 1messageexec was still missingread was still missingAny suggestion where to look?
Hi I have ChatGPT plus and during past week I observed that my weekly limit reset before the supposed end time. For example yesterday I had 25% left for remaining 3 days. This morning it goes back to 100% when remaining 6 days 22 hours. How could this happen?
Hi I have ChatGPT plus and during past week I observed that my weekly limit reset before the supposed end time. For example yesterday I had 25% left for remaining 3 days. This morning it goes back to 100% when remaining 6 days 22 hours. How could this happen?
Hi all:
I've renamed a few devices in Unifi when I swapped some cameras, after which I deleted the Unifi devices and re-added them. I quickly noticed some of device_track entity seems to stored somewhere and linked to old devices in Unifi before renaming. I then deleted Unifi integration and re-added the device that caused the issue, but unfotunately the problem still exists, so somehow the mis-mapping was remembered and there is no way I can change that.
One example is the camera, before renaming I used WiFi to connect, after renaming it uses ethernet, but device_tracker always show the device not_home, dispite that it's connected and shows "online" in Unifi. I have checked the entity in Developer Tool, it maps to the correct ip and mac address of the device. I'm not out of idea and would appreciate if anybody can give me a hand on where to look/try. I have already tried:
Hi 👋 I have a cronjob that simply forward different reports in md format from sub-agents. The cronjob works fine but now it can only send one message to telegram/whatsapp at a time. Is there any way to make it send multiple messages to telegram/whatsapp in one run?
Thanks 🙏
Ok I must confess that I’m really confused by different ways of setting us multiple agent - ChatGPT and a few YT videos made me even more confused.
I understand what are sub-agent and persistent agent, but don’t understand how I could use them.
The idea is simple: I have a project manager that coordinates all agents for delivery, I have a developer, a tester and a security specialist who should always challenge and review risky operations.
I want those roles to have different workspace because they should have different agents and soul MD files. Even more importantly I want them to have different skills. So every document I read suggest I should create persistent agent here.
However I only want to chat with the main agent who delegate work to sub-agent, and I don’t want the sub-agents to talk to each other.
I watched a YT video showing that the main agent can spawn sub-agent based on another persistent agent, is that really what everybody is doing with this setup?
Many thanks
When to use which and why?