u/AlemanCastor

My claw wrote itself a plugin to support Qwen 3.8 27B thinking levels and I share it with you

I run openclaw 2026.6.9 with Qwen 3.8 27B on llama.cpp

As Qwen 3.8 controls it’s thinking levels with the chat template there wasn’t out of the box support on my stack, so my claw wrote this plugin and I asked him to make it public just in case someone else finds it useful.

What it does:
* Propagates thinking levels correctly on the chat template per api call
* optionally it lets you set a different level for heartbeats (I use it with xhigh)
* instruct mode (thinking off) supports changing sampling parameters as recommended by unsloth
* optionally force compaction to instruct mode to prevent timeouts.

Repo is here. Fully written by my claw, I didn’t even look at the code: gitlab.com/moltwithhat/llamacpp-qwen-thinking

reddit.com
u/AlemanCastor — 1 day ago