▲ 17 r/StrixHalo
Am I missing something? Qwen3.8 is very slow on my Strix Halo, while Qwen3.6 27b MTP Q4 can reach 20 tokens even on big contexts, but the 3.8 Q4 is 5 tokens.
reddit.comu/Teslaaforever — 2 days ago
Hi,
​
I have created this for me to see when my hardware does the hard work like when all CPUs/iGPU/NPU works together.
​
apu_top_fancy is the one I'm using now
​
I am sharing it here if anyone is interested in using it