r/hermesagent

Anyone running Qwen3.5-9B locally with Hermes?

Anyone running Qwen3.5-9B locally with Hermes?

I’ve been doing some experiments with a modified Qwen3.5-9B runtime on a 4090. I’m testing a mapped/tiered KV setup where most old context (90%) can sit in normal RAM, while the GPU keeps a small hot working set and only pulls back the regions it actually needs.
I'm still new to Hermes but I've used LM Studio and Ollama a lot.

The main objective is to add massive kv cache runway to local processes.

My Setup / context

  • RTX 4090, 24 GB VRAM
  • Ryzen 9 7900X
  • 64 GB system RAM
  • Windows 11
  • Running local models regularly
  • Current test model: Qwen3.5 9B
  • Running the 9B in bf16 for the current experiments
  • Also already running Qwen3.8 27B Q8_0 in LM Studio, but it is a tight fit on 24 GB VRAM
  • Interested in Hermes as an actual local agent environment, not just one off chat inference
  • Main concern: longer running agent sessions and context growth on consumer hardware

Questions for Qwen3.5-9B / Hermes users

  • Anyone running Qwen3.5 9B locally with Hermes?
  • What hardware are you using?
  • What quant are you using?
  • What context size do you normally run?
  • How does it behave once sessions get genuinely long?
  • Does KV/VRAM pressure become a problem for you?
  • Do you reduce context because of VRAM?
  • Do you notice tokens/sec dropping as context grows?
  • How is Qwen3.5 9B for long running agent sessions?
  • How is its tool use?
  • How is instruction following?
  • How is coding?
  • How well does it retain information from far back in a session?
  • Any failure modes that only show up after long conversations?
  • Anyone compared 9B with the 4B or 27B in Hermes?
  • Anyone using CPU/RAM KV offload already?
  • If so, how bad is the PCIe/latency penalty?

I don't expect anyone to answer all of these questions but any responses are appreciated.

u/Electrical_Offer5667 — 5 hours ago

So peoples with hermes + opencode go subscription, how are you doing past few days?

As you all may know open code go has cut their limits by up to 500%.

Just wondering what people are switching to?

Me - command code + Nvidia nim

reddit.com
u/krrish253 — 9 hours ago

Last tools and cron called on desk

This is still mostly solution in a search for a problem but very satisfying side vibe coded thing. last tools and cron jobs are closest so far to feeling more connected to what Hermes is doing in background

u/rudidit09 — 5 hours ago

One year ago I was a factory mechanic. Today I run a business around my AI agent. all possible because of Hermes

this post is basically a love letter to the Nous team, Peter Steinberg who got me into AI agents, AI in general, and of course you (my dear readers).

I have never talked about my past as openly as I am talking about it today. Only one year back I still worked full-time as an industrial mechanic in a factory. But let's start from the beginning.

I was really bad at school, so the only way to get an apprenticeship was to become a mechanic and start working in a factory. So I did. At first it was pretty good, I even liked it. But towards the end of my apprenticeship I started to feel that this was not what I envisioned for myself and not the job I wanted to do for the rest of my life or raise children with. I developed the deep desire to start my own business (probably a deep insecurity from childhood because I was always the weird kid, the kid that could not get what he wants for himself. I always needed help or somebody else. I also was far from the smartest). Now I believe that if you imagine a specific life for yourself and do not settle for anything else and work hard, you will reach it 100%. That is the faith I have in life.

So I started my own business on the side. Started selling merchandising products for a German car YouTuber. I sold around 15 clothing items over the span of around 8 months. But I had to close that because it did not work out and my partner was not on the same wavelength as me. Lost about 1500€ on that project. After that I tried selling websites because I always liked designing websites. I actually started doing that when I was around 15 years old. I even walked around in industrial areas throwing flyers of my web design business into post boxes. But all did not work.

One day I heard about OpenClaw and from the first minute I knew I needed to have that. I love tracking everything digitally. I have been wearing a Whoop for like a year, also been logging my meals in a food logging app. I even log every cent I spend, even calculate how much gas my car consumes on 100km (around 13L/100km or 18.09 US MPG). I love creating systems and logically working paths and workflows, so having my own AI assistant that would connect to all that information and give me scheduled updates was a no-brainer to me. I tried OpenClaw for a few weeks before switching to Hermes because of the reliability and the promise of it being more cost-efficient.

I had high hopes for my AI agent but let me tell you it completely shattered all expectations. I fucking fell in love with computers, computer systems, websites, AI coding, generating web applications, tinkering with computers and the agent.

This agent has given me the ability to run a YouTube channel, blog, Reddit account, sport, nutrition, business and many other hobbies and interests, all while also studying my first semester of business school. Even though I have never even visited the in Germany normally mandatory Gymnasium to be able to start a bachelor's degree.

What my agent manages for me:

  • tracking nutrition, sport, mental health
  • creating new blog posts out of Reddit posts or YouTube videos. I write the text, Hermes just creates a new post for it and manages the websites and even my VPS server
  • generating asset pictures for videos or thumbnails
  • being my infrastructure advisor about my Hermes agent, VPS, Mac mini and home network
  • being my content creation advisor. He can download full Reddit posts, transcribe YouTube videos with comments and vote numbers, thereby he can tell me what topics are hot and what questions I received most
  • he can give me insight into how much time I spend on learning for my exams, even detailed statistics about time I spent on the specific topic
  • helps me while tinkering and playing around with my tech projects

I'm so thankful for this opportunity. My business is now even centered around Hermes. I literally earn money by doing what I love, tinkering with my Hermes agent, developing new and better ways to set up those agents. And all that on the shoulders of the Nous team, OpenClaw team, various AI teams, and not the least you (the reader).

With AI it is now possible to learn anything just by trying over and over again. I have no friends right now, we all split up one way or another. I'm currently alone with my girlfriend, but she's not really into the technical stuff. With AI and the community of Reddit and YouTube I can learn pretty much anything now. Lately I try to solve every computer problem by coding my way out of it. Something I'm really proud of, and something I have fun doing every day. Thank you Hermes for making this possible.

u/HolmeBengt — 16 hours ago

Not sure what to think about the new Hermes Bot mode

I have a Hermes setup that includes 4 different profiles for family members to use.

One additional profile is for a project I’m working on, in which I was looking to sell AI services to clients. This would involve setting them up with a profile. I later realized that there is some cross-contamination of information and files between profiles. It’s fine for family, but questionable when handling clients’ information. I’m still looking to find a solution for this.

In any case, I’ve seen profiles working generally as silos, each working independently of each other. Now with Bot Mode, I feel the messaging is more like “go ahead and create 10 bots, and they’ll all interact and talk to each other.” That’s a very different approach.

Currently, it seems Bot Mode is more of a change in interface, and exchange between bots/profiles is clunky at best. Over the next weeks, that will surely “improve.”

All this comes at the heels of Grok Bot, which seems pretty cool, but a bit expensive. Hermes saw that it got some traction and decided to copy it, which is fine. It’s just intended for a different use case.

reddit.com
u/RobertoGuerra — 12 hours ago

Kokoro Fast is an excellent local TTS upgrade without much setup

Kokoro Web UI

I recently switched Hermes from Edge TTS to Kokoro-FastAPI running locally in docker and I'm really happy with the results.

For me, the main advantages over Edge are:

  • Better voice quality, with a pretty large selection of voices
  • Voice mixing, so you can blend multiple Kokoro voices together to get the voice you want
  • Very fast local generation — on my RTX 3080 I'm seeing sub-second generation for most of Hermes responses
  • Streaming TTS support - Though there is currently an open bug on this part (see below)
  • No internet dependency for speech generation once everything's setup

Docker setup

I'm running Hermes and Kokoro as separate Docker Compose projects on the same Docker network (webproxy). This lets Hermes access Kokoro directly using Docker DNS at http://kokoro:8880.

My compose.yml for Kokoro looks like this:

name: kokoro

services:
  kokoro:
    image: ghcr.io/remsky/kokoro-fastapi-gpu:latest
    container_name: kokoro
    restart: unless-stopped
    ports:
      - "8880:8880"
    environment:
      - TZ=America/Chicago
      - USE_GPU=true
      - DEVICE_TYPE=cuda
      - API_LOG_LEVEL=INFO
      - ENABLE_WEB_PLAYER=true
      - DOWNLOAD_MODEL=true
    deploy:
      resources:
        limits:
          memory: 3G
          cpus: "1.0"
        reservations:
          devices:
            - driver: nvidia
              device_ids: ["0"]
              capabilities: [gpu]
    networks:
      webproxy:

networks:
  webproxy:
    external: true

I'm using the NVIDIA GPU image here. If you're running CPU-only, AMD, or newer Blackwell hardware, check the Kokoro-FastAPI README because there are different image options.

Also note that in my setup above the webproxy network is managed by Nginx. You may have different networking configuration here.

To start it:

docker compose pull
docker compose up -d

A useful check afterward is:

docker exec kokoro nvidia-smi

This will verify that GPU acceleration is enabled in the container.

Kokoro also has a Web UI at:

http://YOUR-SERVER:8880/web

The Web UI was helpful for testing voices and experimenting with voice mixes. You just add as many voices as you want on the right and change the percentage of each.

Connecting to Hermes

Kokoro-FastAPI exposes an OpenAI-compatible TTS API, so Hermes can use its existing openai TTS provider with Kokoro as the backend.

Unfortunately, some of the settings required for this aren't exposed through Hermes Desktop or Web Dashboard right now (e.g. base_url) so I had to edit config.yaml manually.

The relevant portion is:

tts:
  provider: openai

  openai:
    api_key: local-kokoro
    base_url: http://kokoro:8880/v1
    model: kokoro
    voice: am_michael+am_puck
    speed: 1.2

IMPORTANT: If you aren't using OpenAI APIs anywhere else, Hermes will require api_key to contain a value here. Even though Kokoro doesn't need or use an API key. The value itself doesn't matter.

Obviously, you can change the voice and speed to whatever you prefer, and adding two voices together mixes them:

am_michael+am_puck

I don't know the syntax for mixing with percentages.

After saving config.yaml, restart Hermes.

Because both containers are on the same Docker network, Hermes connects directly to:

http://kokoro:8880/v1

If your containers aren't on the same Docker network, you'll need to adjust that URL appropriately.

A nice sanity check from inside the Hermes container is:

docker exec hermes curl -fsS http://kokoro:8880/v1/audio/voices

If that returns Kokoro's voice list, Docker networking between the two is working.

Current Bug: Hermes Desktop streaming

Kokoro supports input streaming, but input streaming is currently broken through Hermes Desktop's OpenAI TTS path.

There's an open Hermes bug for it here:

https://github.com/NousResearch/hermes-agent/issues/79859

The current behavior is basically:

Hermes generates entire response -> TTS audio begins playback

instead of beginning playback as soon as the first completed sentence/chunk is available. This isn't great, but it's still better than Edge because Kokoro generation is blazing fast and Edge currently has a similar issue. The good news is this is a known issue and is being worked on. Once fixed, Kokoro responses will feel even snappier.

Updating Kokoro

Since I'm using latest as my image, updating Kokoro is just:

docker compose pull
docker compose up -d

That's basically it. I hope someone finds this useful!

reddit.com
u/eXntrc — 8 hours ago

Daily Hermes Agent user — using it for almost everything

I’ve been a daily Hermes Agent user and it’s basically taken over half my life (in a good way).

I’ve been using it for almost everything — managing my server, handling Sonarr/Radarr, and even tackling some pretty complex coding tasks. The costs have been getting a bit crazy since I’ve been running Opus 5, but damn has it been worth it. It’s helped me a lot more than I expected.

Anyone else deep in the Hermes + Opus rabbit hole?

u/fcm — 22 hours ago

I built 16 writing Stylebooks for Hermes because “write better” is a terrible skill

I've been thinking about a problem with writing agents: we keep telling them things like “make this clearer,” “sound human,” or “be concise.”

Those aren't really writing systems. They're vague preferences, so the model eventually drifts back toward its default voice.

I built Agent Stylebooks around a different idea: take editorial systems that real organizations have spent years developing and turn their principles into reusable Agent Skills.

There are now 16 skills, including:

  • google-developer-docs — API docs, tutorials and setup guides
  • govuk — public-service and eligibility content
  • gitlab-docs — concise engineering/product documentation
  • github-docs — product workflows and troubleshooting
  • kubernetes-docs — version-sensitive infrastructure docs
  • mdn-web-docs — web technology explanations and references
  • red-hat-docs — enterprise procedures and runbooks
  • 18f-content — accessible digital public services
  • microsoft-writing-style — product help and support copy
  • mailchimp-content — human customer communication
  • apple-interface-writing — buttons, alerts and UI copy
  • cdc-clear-communication — public-health communication
  • nhs-health-content — patient-facing health information
  • sec-plain-english — financial/investor disclosures
  • w3c-technical-reports — specifications and standards
  • nasa-technical-writing — engineering and technical reports

The useful part for Hermes is that these aren't 16 giant prompts you keep pasting into chat. They're skills you install once and invoke when the artifact calls for one.

For example, the topic could be Kubernetes, but the correct style depends on the job:

$kubernetes-docs for a deployment tutorial.

$govuk For a government page explaining whether companies qualify for an infrastructure program.

$apple-interface-writing For the buttons and errors inside a cluster-management UI.

Same subject. Different writing problem.

You can inspect the catalog first:

npx skills add Neeeophytee/agent-stylebooks --list

Install one:

npx skills add Neeeophytee/agent-stylebooks --skill google-developer-docs

Or install the whole collection:

npx skills add Neeeophytee/agent-stylebooks

Then the actual instruction to Hermes can stay simple:

Use $google-developer-docs to turn these verified implementation notes into a setup guide.

The skills are also deliberately conservative about substance: changing style must not change facts, commands, versions, legal meaning or technical behavior. The repo includes source/provenance notes rather than just scraping style guides into prompt files.

It's free/open source(MIT):

Repo: https://github.com/Neeeophytee/agent-stylebooks

I'm especially curious how people here would use this with Hermes' own skill system. Which writing system is missing that you'd actually install?

u/ShilpaMitra — 12 hours ago

Should i use hermes bots or telegram bots?

Hi, I'm experimenting with Hermes, and so far I like it; it fits my needs. I've already set up 3 Telegram bots linked to dedicated profiles. Is it worth it to change to the new Hermes bots functionality?

I don't really understand how it's different or if it has different benefits.

reddit.com
u/F1nch74 — 11 hours ago
▲ 3 r/hermesagent+1 crossposts

Poor "intelligence" with Gemma4

I am testing Hermes with Gemma 4 (gemma-4-26B-A4B-it-OptiQ-4bit running with OMLX) on a Mac mini M4 64GB:

it is fast in terms of response times, but it sucks IMHO in terms of "intelligence" and "agency".

After several days spent (with the help of Claude) debugging and prompting I gave up: even a simple daily cron task (message me the output of curl -s 'https://wttr.in/{city1,city2,city3}?format=3') fails with hallucinations.

Same for a briefing (extracting news headlines from RSS feed links, for which I provided the python code as it was unable to code a working version of it): some times it works, most of the times it doesn't.

I was having much better results with much earlier explorations of OpenClaw.

Where am I wrong?

reddit.com
u/mkeee2015 — 20 hours ago

Sam Altman on Personal AI. Anyone already doing that using Hemres?

Sam Altman envisions a shift toward "always-on" personal AI assistants that maintain continuous context across your digital life by monitoring screens, meetings, and communications, all could become a reality within six months.

Is anyone already doing this in the truest sense? like Hemres monitoring your screen and learning about you and also doing some browsing work on its own? Hermes monitoring your other devices and connecting to you through proactive voice updates and screens?

I am still new to Hermes, using it for less than a week. I am looking for comments which can expand my thinking on what Hermes can do and what smarty people here are already using it for! Please enlighten me.

When I first read this comment of Sam Altman, my first thought was, some smart people here might already be having this setup using Hermes. Or am I overestimating what can be done using Hermes? Pardon me if this is very naive or stupid question, I am new to Hermes and AI in general.

Looking forward to hear from you all!

reddit.com
u/Gloomy-Recover-9702 — 23 hours ago

Nous Research just casually added 900k context window

Just a friendly post to let you all know that Nous Research announced increased available context window on GPT 5.6 to 900k on Discord the other day.

So if you happen to use GPT 5.4 or 5.6 (does not work with 5.5) you can now custom set your context window up from 270k to 900k.

Edit: 24 hrs of running seems to reveal some coherence issues. Logically it should improve the output of bigger tasks, but the output I received from this run has been below expectations. Mind you the issues might be caused by something else entirely. Will investigate further tomorrow and update accordingly.

reddit.com
u/galimatis — 21 hours ago

anyone know of a great uncensored llm to host even locally?

not for anything crazy but these stupid commercial llms cant even work on security for one of my sites im building- its so annoying bro

reddit.com
u/RosalbaaaaAAbbey — 19 hours ago

Using Hermes with 5.6 Luna

After the price increases, I switched the backend from DeepSeek flash to 5.6 Luna just to test the waters. I keep hitting TPM errors even in a brand new session for the simplest of tasks. What’s your take on this? What could be the issue?

reddit.com
u/stochastic-36 — 18 hours ago

Hermes models

I been using Hermes and running ChatGPT models inside it. However, I’m burning through my limits way too fast(the pro subscription), especially for lighter daily tasks where I don't need a heavy model.
For those of you primarily using Hermes, what cheaper or lighter models are you pairing with it for everyday tasks to save your tokens.

I tried local models like qwin 3.6 but it wasn’t very helpful.

reddit.com
u/Frosty-Mixture-3513 — 1 day ago
▲ 5 r/hermesagent+1 crossposts

I want to quit using AI cause the cost is just too high

I had a GitHub CoPilot Pro license for a long time for $10 per month and this worked fine. Then all of a sudden the price went up in April and I had to make a new plan.

I then tried ChapGPT Plus for month for $20 per month, which was more than I really wanted to spend, but it's all I could find at that point. I was using it mainly for writing code for my personal side projects and some work projects (but I'm not a software developer at work).

I found Opencode Go and that sort of worked for a while even though I had to use OpenCode Zen with a $10 credit limit. Around this time I found Hermes and started implementing it with my work and personal stuff.

Now DeepSeek came out with new pricing and forced Opencode's hand and now I'm back to where I was with GitHub CoPilot, but a lot more dependant on AI because now I use Hermes too. And I feel like it took some time to get Hermes to the point where it is now.

I feel like we'll keep going through this cycle until we eventually realise that AI costs are too high or you just make peace with it and pay $20+ per month.

I don't spend a lot of money on subscriptions. I have Netflix and Google One in addition to OpenCode Go.

So what now? Quit? Use Mimo. Make peace and go with ChatGPT Plus until I find something else? I don't want to try Command Code because it just sounds like another Opencode Go and I'll be back here in 2 months.

reddit.com
u/lkn-ant — 1 day ago

What a session costs in Hermes per model

GPT-5.6 Sol Pro costs 504× more per session than MiMo-V2.5 on Hermes.

u/maferase — 1 day ago

Hermes cant do any accurate Animations/Rendering

Hermes (Using DeepSeek V4 Flash) cant for the life of it render anything accurately, is there any Skill or anything I can do to improve its Capabilities? Examples of what I mean below:

prompt: can you mark the input and output (image attached)

it researched how such a drive looks like and all about it: ca. prompt: Research how a three-ring-drive works and replicate it in an animated inline-rendered Visual (

reddit.com
u/someoneyouknow23 — 1 day ago

I tested new Hermes Bot Mode plugin here’s what happened

I tested Hermes Agent’s new Bot Mode with 3 AI agents here’s what actually worked (and what didn’t)

Hermes Agent recently released Bot Mode, and I wanted to test whether it actually makes multi-agent collaboration easier or if it’s mostly a nicer UI for existing Hermes profiles.

So I built a small 3-agent Stocks research team:

🔎 Researcher — gathers evidence from public filings and market data
⚠️ Risk Analyst — challenges the numbers, assumptions, unsupported claims, and missing risks
📝 Thesis Editor — turns the reviewed evidence into a balanced final report

I only gave the task to the Researcher.

The Researcher gathered the evidence and handed it to the Risk Analyst. The Risk Analyst reviewed it and was supposed to hand the result to the Thesis Editor.

So instead of:

Human → Agent A → Human → Agent B → Human → Agent C

I was testing:

Human → Agent A → Agent B → Agent C → Human

And for the most part, it worked.

What I like about Bot Mode is that the agents aren’t just invisible calls in a workflow. Each bot has its own identity, role, tools, skills, memory, and persistent conversation, so you can actually inspect what each specialist is doing.

But there are some important limitations.

It’s also not a full workflow/DAG engine, and the handoffs don’t mean you’re getting guaranteed parallel execution.

My current mental model is:

🧠 Hermes Profiles = specialized isolated brains
👥 Bot Mode = gives those brains identities, persistent rooms, and communication
🗂️ Hermes Kanban = better when you need structured tasks, dependencies, and project orchestration

For complex multi-agent projects, I’d still lean toward Kanban. But for maintaining a group of persistent AI specialists that can talk and delegate to each other, Bot Mode makes the experience much easier to operate.

I recorded the entire experiment, including creating the bots, configuring them, the handoffs, the failed delegation, scheduled jobs, and the differences between Bot Mode, Profiles, and Kanban:

🎥 https://youtu.be/w3VI6zC4_0I

I’m curious how others building multi-agent systems think about this approach.

u/viky_shetye — 1 day ago