u/rocket4time

▲ 3 r/Qwen_AI+1 crossposts

Qwen3.5-9B / Qwen3.8-28B quantization for 11GB VRAM

I have an 11GB VRAM limit and I'm trying to choose the best quantization for a local LLM.

I'm considering:

  • Qwen3.5-9B Q8
  • Qwen3.8 -27B Q2_XS

I initially looked at very aggressive quants like IQ2_XS / Q2, but I'm worried about the quality degradation.

What would you recommend for the best quality/VRAM trade-off within 11GB?

I'm particularly interested in real benchmark results rather than just perplexity. The model will be used for factual QA, RAG/grounded answers, and general reasoning.

Has anyone compared these quants on actual tasks? How much quality do you lose going from Q5/Q4 to Q3 or Q2?

What would you pick with an 11GB VRAM hard limit?

reddit.com
u/rocket4time — 5 days ago