Qwen3.8 27B effort levels
I do not understand how to get the model to limit reasoning. I am using unsloths q8 and q8_xl version in unsloth desktop. Combinations i have tried:
Low effort with a detailed prompt = 10+ minutes of thinking.
Low effort with a simple prompt = 10+ minutes of thinking.
Low effort with a simple prompt that requests a rapid prototype = 10 + minutes of thinking.
medium effort with simple prompt = 10+ minutes of thinking.
It seems like it always uses xhigh no matter what i do. However on random occasions it has thought for around 1 minute but i can't reproduce it.
Same issue when i connected it to hermes agent, and also happens on lmstudio but i dont even have effort level options in lmstudio so that is expected that it would default to xhigh.
system: windows 11, 1x tesla v100 16gb, 1x tesla v100 32gb, 32gb of ram, ryzen 3800x.
using mtp and originally used tensor parallelism but it would randomly give me issues and the api wouldnt respond so i disabled it. I get 56 tps with it off so you can judge "10 + minutes of thinking" appropriately.