Qwen3.8 Q4_k_m 1M context Strix Halo / 3080 -> 45 tps
Been playing qwen3.8 for a couple of days, on a AMD 128GB strix machine, with oculink eGPU (3080ti) 12GB.
Running Q6, 262K context
on just strix halo without MTP 10tps, with MTP with n=4 24 tps, with FastMTP offloaded to 3080 with n=4 about 28 tps
Q4_k_m 262K context
Splitting layers 5GB weights & FastMPT on 3080 and rest on iGPU~ 53 tpsQ4_k_m, 1M context, k-q8, v-q4
Splitting layers 2GB weights & FastMPT on 3080 and rest on iGPU~ 45 tps
Still isn’t as good as a 5090 but AMD still show 60-70GB free so I can still load qwen3 embedding and a qwen3 Reranker to go with everything
Now need to see how good this model really is compared to qwen3.6-35b that I have running on 4x3090 and runs at 450 tps….