M4 Pro 48GB - Qwen3.8
So unfortunately I’ve tried the basic 27B q6 mlx from mlx-community and it only does like 8t/s. This model overthinks a lot so that mixed with 8t/s makes a simple task seem like days.
Any ideas how to speed it up? Or did anyone test q4 vs q5 vs q6?