▲ 85 r/LocalLLM
Are q4 quants suddenly OK now?
Now that the first wave of Qwen3.8 "muh benchmarks" is coming to a conclusion, can we share some actual real-world notes on results from different quantizations?
I avoid q4 based on my experiences with all earlier Qwen models, too many loops and inaccurate results in my agentic usage.
Is q4 suddenly usable? I see people benching that it's not that far from q8, and I'd love to claim back a bit of context and concurrent from my vram if so.
u/bigb159 — 1 day ago