u/MrWidmoreHK

Qwen 3.8 27B running at up to 36tps on the Halo

Qwen 3.8 27B running at up to 36tps on the Halo

Opus 4.6 level intelligence at home for les than $3k

Got Qwen 3.8 27B running at up to 36 tps on the AMD Strix Halo

⚡ ROCmFP4 block quant (13.5 GB)

⚡ MTP Speculative Decoding (2.9× speedup)

⚡ Full 262K context in ~33 GB RAM

https://github.com/julianmb/q38rocm

u/MrWidmoreHK — 5 days ago