▲ 26 r/Qwen_AI

Would you agree me that we need a low para MoE model like Qwen3.8 14B A2B

I have only 16GB (exactly 15,6GB) free completely for LLMs, which allows me to run Gemma 4 12B, Ornith 1.0 9B or Qwen 3.5 9B.

The problem: 9B and 12B models are normally dense models which is why they dont run that fast on a 128bit bus.

GPT OSS 20B (MoE) runs fast and the Model fits in my RAM, but with Context it dosent.

Solution: If Alibaba would release a 14B MoE model (like AMD's 16B MoE), it could fit in any 16GB VRAM + Context + fast decode speeds.

I think that most of the people want a Qwen 3.8 100B~ MoE model but I and other 16GB users cant even do something with a 27B model.

reddit.com
u/Appropriate_Lead439 — 1 day ago
▲ 34 r/LocalLLM+1 crossposts

New FastFlowLM v0.9.46 and FLM joins ROCm organization.

- FastFlowLM is offically maintained by AMD.
- New NPU-ROCm channel on AMD Developer Community (Discord).
- Modelscope support.
- More control over Qwen-VL models.
- ~10% faster decode and prefill for Qwen 3.6 35B A3B .

https://preview.redd.it/h8pcx3az62gh1.png?width=558&format=png&auto=webp&s=9990ca5b37e6e131cd6bed03ffe7ce3c76ae45da

Releases · FastFlowLM/FastFlowLM

ROCm/FastFlowLM: FLM app mirror

reddit.com
u/Appropriate_Lead439 — 22 days ago

Auto SR option gone and package bugged should get fixed

The Microsoft dev team has officially acknowledged, that the Auto SR option is gone and the Auto Super Resolution package is bugged, and "they will look into it (Wir prüfen dies)".

You can upvote this report, so that Microsoft will perfer this fix compared to other problems.
Feedback post: https://aka.ms/AA11rirt
I would expect it to be fully fixed at the end of the month.

Even tho I know how buggy and limting Auto SR was, it's still in a preview state and a "nice to have" feature.

Is the Auto SR UI also gone for you?

u/Appropriate_Lead439 — 2 months ago
▲ 5 r/AMDLaptops+1 crossposts

Why Auto SR underperforms on AMD based devices compared to ARM Snapdragon X Elite?

Hey everyone,

I've been doing some deep-dive testing on the Ryzen AI 7 350 /w Radeon 860M, LPDDR5X 7500MT/s, XDNA 2 NPU to understand how NPUs impact gaming performance, specifically regarding Auto SR.

On ARM-based Copilot+ devices (Snapdragon X Elite), Auto SR works beautifully. Offloading the upscaling from the GPU to the NPU allows 720p upscaled to 1080p to hit the exact same FPS as native 720p.

However, on AMD x86 architecture, the story is completely different. Testing native 720p vs. Auto SR 1080p resulted in different framerates. In fact, running 720p upscaled via the NPU dropped performance so much that it matched native 1080p FPS, defeating the whole purpose of upscaling.

Since the Auto SR option wasn't available temporarily, I simulated the NPU gaming workload by running Gemma 4 E2B on the NPU via Lemonade while benchmarking Minecraft (Low Settings 4K).

Here is the data I gathered across different TDP steps using HWiNFO64:

Scenario TDP FPS TPS (for LLM) GB/s (RAM=
Gaming Only 28 Watts 53 FPS 37GB/s
Gaming Only 15 Watts 45 FPS 29GB/s
Gaming and LLM with NPU 28 Watts 45 FPS 14.5 TPS 52GB/s
Gaming and LLM with NPU 15 Watts 36 FPS 13.5 TPS 49GB/s
LLM with NPU Only 28 Watts 18 TPS 30GB/s
LLM with NPU Only 15 Watts 18 TPS 30GB/s

(AI helped creating this table by my real data and inputs)

Analyse and conclusion:

- Activating the NPU at 28W drops the gaming performance to exactly 45 FPS, which is identical to running the GPU alone at 15W.

- Even though memory bandwidth scales up to 52 GB/s when both are active, it is not maxed out with the theoretical bandwith of 128GB/s.

- Arm based devices dont get this big of a performace hit by Auto SR.

What are your thoughts on this? Is this a problem that can be fixed via software, or are we looking at a hardware limitation?

reddit.com
u/Appropriate_Lead439 — 2 months ago

Microsoft has about 3 days left to release Auto SR and I think it should be released on Tuesday, because there will be an optional update available, which is in my opinion, the last chance to launch Auto SR. The last 3 Insider Updates were on a Friday, but the last April day is a Thursday, which is why there is only the 28 of April left for me.

reddit.com
u/Appropriate_Lead439 — 4 months ago