Qwen3.5-9B / Qwen3.8-28B quantization for 11GB VRAM
I have an 11GB VRAM limit and I'm trying to choose the best quantization for a local LLM.
I'm considering:
- Qwen3.5-9B Q8
- Qwen3.8 -27B Q2_XS
I initially looked at very aggressive quants like IQ2_XS / Q2, but I'm worried about the quality degradation.
What would you recommend for the best quality/VRAM trade-off within 11GB?
I'm particularly interested in real benchmark results rather than just perplexity. The model will be used for factual QA, RAG/grounded answers, and general reasoning.
Has anyone compared these quants on actual tasks? How much quality do you lose going from Q5/Q4 to Q3 or Q2?
What would you pick with an 11GB VRAM hard limit?
u/rocket4time — 5 days ago