Would you agree me that we need a low para MoE model like Qwen3.8 14B A2B
I have only 16GB (exactly 15,6GB) free completely for LLMs, which allows me to run Gemma 4 12B, Ornith 1.0 9B or Qwen 3.5 9B.
The problem: 9B and 12B models are normally dense models which is why they dont run that fast on a 128bit bus.
GPT OSS 20B (MoE) runs fast and the Model fits in my RAM, but with Context it dosent.
Solution: If Alibaba would release a 14B MoE model (like AMD's 16B MoE), it could fit in any 16GB VRAM + Context + fast decode speeds.
I think that most of the people want a Qwen 3.8 100B~ MoE model but I and other 16GB users cant even do something with a 27B model.