How do I set thinking level with harnesses like Pi or Opencode?
Qwen 3.8 27b is known to be overthinking a lot as the default setting is XHigh. How do I set a lower level in harnesses like Pi or OpenCode? Thanks!
Qwen 3.8 27b is known to be overthinking a lot as the default setting is XHigh. How do I set a lower level in harnesses like Pi or OpenCode? Thanks!
Hi! Anyone got VLM MTP working on Gemma 4 (e4b or 26B) in oMLX?
I've been going round in circles. The docs say Gemma 4 uses a separate assistant model for speculative decoding via VLM MTP — enable `vlm_mtp_enabled`, point `vlm_mtp_draft_model` at the assistant checkpoint. Sounds straightforward.
**What I've tried:**
- e4b 4-bit + `gemma-4-E4B-it-qat-assistant-bf16`
- 26B OptiQ + `gemma-4-26B-A4B-it-assistant-bf16`
- 26B Unsloth QAT oQ4 + `gemma-4-26B-A4B-it-qat-assistant-bf16`
- 26B MLX community variant recommended by Google documentation
All on M4 Max 36GB (not bumping into memory limits with this), oMLX 0.5.2. VLM MTP enabled, correct assistant model selected.
**Results:** consistent slowdowns across the board. e4b went from ~80 to ~75 tok/s, 26B from ~68 to ~66 (OptiQ) or ~83 to ~74 (QAT). Assistant model loads internally but adds overhead with zero benefit. No acceptance stats exposed anywhere.
So... am I missing a knob? `vlm_mtp_draft_block_size`? Some other setting not exposed in the admin UI? Or is VLM MTP just not ready for Gemma 4 on Apple Silicon yet?
Would love to hear from anyone who's actually seen a speedup on e4b or 26B.
Thanks!
Hi, I love Hermes dearly when it works BUT it's a constant struggle. I'm on M4 Max Studio with Tahoe, usually running Deepseek 4 Flash. Hermes installed via the official curl command. Sometimes Hermes Desktop but mostly Hermes WebUI and Hermex on mobile.
- every day, something breaks (Telegram stops responding, gateway silently dying, some cronjobs silently fail, today Agentmail MCP weirdly starts showing 0 tools and stops working...). Hermes doctor never finds any issue.
- every day I get several permission prompts for Python, Node, sometimes Microphone, Apple Music, Photos... Until I manually approve it on the Mac, work stops. Given that I usually use Hermes remotely from a phone, laptop or tablet, this is a hard blocker (yes, I can Remote desktop and approve, but that defeats the purpose of an autonomous systm).
I'm not a developer, but not a total IT noob either. I use Hermes for assisting with video production (various reports from timeline and footage, automating transcriptions via Macwhisper, searching for and downloading footage for documentaries...) and with my homelab (Proxmox, Synology, Home Assistant, all sorts of stuff).
As such, for my work-related stuff, I need reliable access to my Mac filesystem, so running it directly on the system is the best option. When it works, it's absolutely stellar.
I can run Hermes inside a LXC on my Proxmox, thus removing the need for repeated MacOS perm popups, but that has other downsides, such as flaky computer use (with remote cua-driver, unreliable), not being able to directly point to paths (Mac filesystem IS mounted into the LXC but paths differ)...
Is there any solution to this mess? I do understand some of this is a user error, and I can't expect polish of a commercial product, but given the fact people buy Mac Minis to run Hermes and the entire point of Hermes is autonomously assisting the user in the background, this can't be the intended state of things. It needs constant handholding.
Thanks!
EDIT: I move the complete instance over to a LXC and mapped the necessary paths from Mac into it. The paths are 1:1 identical, and it seems to work. No more TCC prompt requests. Even cua-driver for computer_use works remotely. Lets see if this sticks. Leaving this here for posterity.
I'd like Hermes to auto-updare itself and Hermes Webui but the nightly cron job it has set up fails repeatedly, with seemingly different culprits (last time the gateway broke with Error during OpenAI-compatible API call #4: cannot import name '_plan_tool_batch_segments'
from 'agent.tool_dispatch_helpers' (/Users/stooovie/.hermes/hermes-agent/agent/tool_dispatch_helpers.py))
What's yall's update strategy? I thought for an autonomous agent, this should be THE number one thing it does well but apparently not. Running DS4 Flash or Hy3.
Thanks!
Hermes installed via official command:
curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash
After updates, stuff (python, node, Hermes, the Hermes Desktop app...) randomly asks for permissions (file, photos, microphone...) again. I mostly use it remotely from tablet or phone, so I don't see the perm request popus, and so it fails silently.
This defeats the purpose of an autonomous agent. It's NT just the GUI desktop app, but also the cua-driver and python3 installed by Hermes.
Any way around this? Gatekeeper allows installations from anywhere (probably unrelated). Tahoe 26.5.2 on M4 Max Studio.
I'm constantly fighting with Hermes' internal content redaction. It even redacts text from websites scraped by crawl4ai, making it useless. Basically anything written by write_file is considered as private content and gets redacted. Read_file does it too.
Any tips? Thanks!
I have Mnemosyne up and running for a week, and the recall ability is as bad as the default system. It just doesn't recall without handholding and reminding.
Mnemosyne is confirmed active. Nous' default setup by terminal script on an M4 Mac, and also in a Proxmox LXC via Proxmox Helper Scripts, same behavior. Everything is fully updated.
Llm-wiki doesn't help either (not sure if it's supposed to)
I tried Holographic, had about the same experience. Using DS4 Flash by default - maybe that's the culprit?
Any tips? I did search here, most comments say Mnemosyne fixed the memory issues :) Thanks!
EDIT: asked Hermes itself about it and it's telling me that the mnemosyne integration has "auto_sleep"=false by default, which is in direct contradiction with Mnemosyne documentation. With False, memories never consolidate and do not create the vector database. Very weird, and probably a bug in Hermes. I'll see if this fixes the issue somewhat.
EDIT2: consolidated, fixe config, all tested... And the very first thing in a new session, it completely blanks out again. And AGAIN in the next. Explicit system prompts stating any mentions of github or tokens MUST look up a Memory or wiki page is ignored. That looks like a big Hermes issue to me.
Here are some of my use cases of a non-IT, non-developer Hermes user (I'm a broadcast video editor), maybe as an inspiration :)
Over the last two weeks, I have used Hermes Agent for
Obviously I recommend having backups or snapshots for this.
- turning off indexing on my old Synology that is useless for my use, and can't be disabled from GUI - much quieter now
All this stuff saved like a day of menial work over the last week alone - no exaggeration. And that's just getting started.
I haven't yet set up many cron loops as my needs are rather scattershot and one-off, but I'm sure I'll find out some later.
I'm using a combo of Deepseek 4 Flash and locally running GPT OSS 20b (via oMLX) for the easier stuff. Runs incredibly well.
Zen 1.21b3 on Mac. Cannot add Essentials in any way. I have 8 existing, 12 is max.
- Enable Container-specific Essentials is OFF
- unpin - Add as Essentials still not possible.
I see a suggestion to disable Containers which are apparently enabled by default, but Zen says it would close 662 Container tabs, which I have absolutely no idea what that means. I don't consciously use Containers.
Tips? Thanks!
Hermes 0.17 is now telling me it can't analyze local images:
****
`vision_analyze` can’t read a local file path directly because it delegates to a remote service.
Use a data‑URL (base‑64) or serve the file over HTTP, and you’ll get the same detailed description you’d expect.
****
(the bytecode won't fit the context and it fails)
No cloud providers at all, GPT OSS 20b main with Gemma 4 e2b for vision. Working flawlessly in OWUI but refusing in Hermes.
Any tips? Thanks!
SOLVED: it's an oMLX bug (a regression). No models other than Gemma 4 26b properly load as VLM.
Is it me or is there no way to answer or approve Hermes' questions and approval requests? Like /new asks if it can actually create new session, but there are no buttons and answering literally (how it asks me to) with "!approve" doesn't work.
Using Element X as recommended per Hermes documentation. I see no mentions of this there. Reactions do not work either. There's literally no... reaction.
It also answers a LOT slower than Telegram. Same question, same Hermes instance, same model.
I'm confused (and my setup apparently as well) about Hermes Desktop and Hermes in CLI. I have installed Hermes with the Desktop app from https://hermes-agent.nousresearch.com/ and nothing else.
Are the Desktop and CLI supposed to be two separate entities with separate settings? Right now, settings from one aren't honored in the other
I have set up all sorts of things like my SearXng instance for web search, and it did work properly but now that's gone. Hermes CLI doesn't see that config at all (and searches the web by brute force with chrome headless), Desktop does but doesn't use it.
Tried to set it up again with "hermes tools" but the web search setup wizard now only offered me Nous sub and Firecrawl. There were many other options a week ago when I set this thing up.
Same with Telegram. I have set it up, it did work, then it stopped working. In both Dashboard and Desktop, the toggle was ON, and couldn't be switched off (the toggle would just re-enable after 2s). hermes setup did let me reconfigure with new bot and token but when I checked the actual .env, the old one was there. It only began working after manually changing the env token variable.
I'm going slightly mad here. It all worked yesterday, all broken today. Any tips? Thanks!
tl;dr toggle in GUI doesn't save, config.yaml is ignored
v0.16.0 81eaedd, via Hermes Desktop app. I have set up Telegram and Home Assistant. Both worked fine. Since I have disabled Telegram.
Today, I can't enable it and disable Home Assistant. Toggles toggle right back. No further updates available. Anyone else?
~/.hermes/config.yaml has Telegram enabled True
Tried making a new Telegram bot and reconfigured gateway (and restarted it). Doesn't work. New token is apparently not in ~/.hermes/config.yaml (the old one is). hermes doctor doesn't see issues.
I use all sorts of interfaces to Hermes Agent depending on context, goal and device (Hermex app on iPhone, OWUI, cli, Hermes WebUI...) I'd like to have a unified chats list with all of these, but it's a bigger problem than I thought. Some I can see only from CLI, some show up as "API chat" only...
I understand why Hermes chats do not populate in OWUI but the other ways, when I talk directly to Hermes, I don't understand.
Any tips? Thanks!
Using LM Studio's own Locally AI app breaks the JIT eviction system - when you switch models in the app, they get added on top of the already existing ones, until total RAM exhaustion.
Just a reminder if someone else's having this issue. Filed at Github.
There's no way to set up an API key with a custom endpoint. I'm running oMLX on Mac (so, OpenAI endpoints) but the Desktop app refuses to read models from it. It tells me to set up an api key (which I do have), but there's no way to actually enter it.
Any tips?
I may be doing something wrong but I'm not overly impressed by Gemma 4 12b from yesterday. Compared to 26b, it runs as 1/3 of the speed (70t/s vs 25 on M4 Max Studio) and really sucks at non-English languages. I'm using the gemma-4-12B-it-mxfp4 quant from mlx-community (the 26b is the same quant). It's said to have MTP but omlx says otherwise.
Also it's leaking <audio> tags into text, but that could be an omlx issue.
Any tips or comments?