
So. about the speed of Qwen 3.8 27B Q2_K_XL on 3080 12GB
this screenshot is without MTP usage , fully offloaded onto the GPU. using LM STUDIO
i wanted to ask you guys if theres a way to make it even faster , as MTP really didnt help and is infact slower due to vram overflow and that im on 16GB DDR4 which is disgustingly slow to load models on so i depend on my gpu for every model.
this is unsloth's GGUF quant