▲ 23 r/AIProgrammingHardware+1 crossposts

Qwen3.8-27B Q8 MTP benchmarks on Strix Halo — MTP is actually making it slower. Are others seeing the same?

I've been testing Qwen3.8-27B Q8 on a Ryzen AI Max+ 395 / Strix Halo system through Lemonade + llama.cpp, specifically to see whether the new MTP speculative decoding support actually improves generation speed.

I kept the prompt and output length identical between runs:

  • 96 input tokens
  • 1024 output tokens
  • Temperature 0
  • Same Qwen3.8-27B model/quant
  • Flash Attention enabled
  • --no-mmap
  • Only backend/MTP settings changed

These are the results so far:

Backend MTP setting Generation speed TTFT
Vulkan Off 9.159 tok/s 0.758 s
Vulkan n-max=1 6.579 tok/s 0.950 s
Vulkan n-max=3 7.122 tok/s 0.764 s
ROCm Off 6.534 tok/s 0.715 s
ROCm n-max=3 4.689 tok/s 0.688 s

So on my machine:

  • Vulkan + MTP n=3 is about 22% slower than Vulkan without MTP.
  • Vulkan + MTP n=1 is about 28% slower.
  • ROCm itself is about 29% slower than Vulkan without MTP.
  • ROCm + MTP n=3 drops another ~28% versus ROCm without MTP.
  • Overall, Vulkan without MTP is almost 2x the generation throughput of ROCm + MTP in this test.

For MTP I'm loading llama.cpp with:

--spec-type draft-mtp --spec-draft-n-max 3

(and also tested n-max=1 on Vulkan).

For ROCm, I'm using Lemonade's current stable ROCm backend. Lemonade reports the llama.cpp backend as b10397; the bundled ROCm/TheRock stack appears to be ROCm 7.13.x. I haven't tested ROCm 7.14 yet.

It's surprising to see that with MTP there's a pretty substantial regression on both Vulkan and ROCm.

I'd be interested to compare with other Strix Halo owners:

  1. Are you seeing MTP actually improve Qwen3.8 throughput?
  2. What --spec-draft-n-max value works best for you?
  3. Are you using Vulkan or ROCm?
  4. Which ROCm version / llama.cpp build?
  5. Does ROCm 7.14 materially improve Strix Halo performance versus 7.13?
  6. What Qwen3.8-27B quant are you using?
  7. If you're getting a significant MTP speedup, what kind of draft acceptance rate are you seeing?

I'm mainly trying to figure out whether these numbers are normal for the current llama.cpp MTP implementation on Strix Halo, or whether something is wrong with my setup.

At least with my current stack, Vulkan with MTP disabled is very clearly the fastest configuration I've tested.

reddit.com
u/SecuredStealth — 3 days ago

How do you get a family reunification appointment today?

Hi, I wanted to know if the cluster fuck of family reunification appointments has made any progress towards getting resolved. What is the current way to secure an appointment with AIMA today?

Thank you

reddit.com
u/SecuredStealth — 8 days ago

Anyone noticed the last few rocm nightlies completely broke inference?

I’ve been updating the rocm nightlies since 1 week trying to find one which won’t break inference. Currently all the recent ones break the models because the models are just generating garbage like /////// and so on… anyone having these issues? I’ve switched to Vulkan for now

reddit.com
u/SecuredStealth — 14 days ago

What are the latest rocm enhancements I should be running?

I recently setup lemonade on my strix halo, and pointed it to the rocm nightly builds already. But, over a few sub reddits and different posts, I see that there are various other "better" forks of rocm for the strix halo. So, I wanted to know, that as of today, what are the recommended enhancements for rocm that I should have for maximum performance?

thanks

reddit.com
u/SecuredStealth — 21 days ago

Griffith Observatory Planetarium tickets - is it difficult to buy tickets?

Hi everyone,

I'm planning to visit Griffith Observatory on a Friday and would like to attend the 7:45 PM "Centered in the Universe" planetarium show.

According to the Observatory website, tickets for that show only go on sale at 7:00 PM and cannot be purchased online in advance.

Does anyone know how the ticketing actually works in practice?

Can we start lining up near the box office or ticket machines at around 6:30 - 6:40 PM, or is there no official queue until 7:00 PM? Since there are apparently multiple ticket machines, where is the best place to wait to maximize our chances of getting tickets?

Also, does this particular Friday evening show commonly sell out immediately?

Recent firsthand experiences would be greatly appreciated. Thanks!

reddit.com
u/SecuredStealth — 25 days ago

How do you find "recommended" skills?

There are about 90K skills on SkillsHub, what's the right approach to finding "recommended" and must use skills? yes, I know it varies from every use case but what would be the recommended common skills apart from the built in ones. How do you find them? There's no rating as such for the skills...

reddit.com
u/SecuredStealth — 1 month ago
▲ 2 r/Rag

AnythingLLM RAG: Getting very inaccurate answers

Hi, I'm running AnythingLLM on my Strix Halo 128GB, loaded with GPT OSS 120B. I've set the context size to around 128K tokens. I'm loading a bunch of markdown files - directly into the chat - so context stuffing gets activated, but, I'm getting very very inaccurate answers on my files. They are all markdown files. What do I need to do to get accurate answers? I've tried it in RAG mode as well, but that didn't help.

reddit.com
u/SecuredStealth — 1 month ago
▲ 0 r/krakow

Silly question: Where can I buy Alpen Muesli in Krakow?

So… yeah I couldn’t find it on store shelves. Do I have to buy it online? Is it available locally somewhere? Thanks

reddit.com
u/SecuredStealth — 2 months ago

Webcam stutter on Caldigit TS4

Hi, I'm using a TS4 dock connected to my 16 inch M3 Macbook Pro. To the dock is a single 49 inch display connected via DisplayPort outputting 5120x1440 @ 240. There's also wired ethernet, speakers, a keyboard and a mouse connected to it. And, also a webcam. I've noticed that the webcam ocassionally stutters during calls. I've shifted the connection to various different ports on the dock, updated the dock to the latest firmware as of writing this, tried reducing resolution of the webcam, but none of this helps. I'm using the official included thunderbolt cable for connection. Are there any other steps I can do to fix this? FWIW, the issue appears of 2 different brands of webcam.

Removing almost all of the connections and trying hasn't helped either. I've also tried reducing the display resolution, but that hasn't helped. I'm using the official cable to connect from the webcam to the dock.

Are there any other troubleshooting steps?

reddit.com
u/SecuredStealth — 2 months ago

Hi, I'm newbie to simracing, I haven't ever done in the past. I'm considering purchasing the RS50, apart from the cost, is there any reason why this might be a bad initial purchase? I don't want to start with the basic Logitech steering wheels only to upgrade again. So again, keeping the cost apart, please let me know if there's any reason why I shouldn't start off with the RS50?

reddit.com
u/SecuredStealth — 4 months ago
▲ 13 r/krakow

Hi, I’ve purchased a king size bed from IKEA which I don’t like. This was delivered to my home. But, how do I return it back to them? I don’t have a car and IKEA doesn’t appear to have pick ups. Any suggestions?

reddit.com
u/SecuredStealth — 4 months ago