u/0rand

▲ 36 r/oMLX

New OMLX with NPU released! Qwen 3.8 27b MTP on M5 Max - results

Just fresh from the presses! oMLX 0.6.1 build 2323 with NPU support.

Checkpoint: Qwen3.8-27B-oQ4e-fp16-mtp (scottlowry/Qwen3.6-27B-oQ4e-fp16-mtp)

Lightning MTP - on

Qwen ANE prefill - on

ANE Prompt Block 2048

MLP on ANE 53%

MLP Layer limit - 64

Use both ANE - off

Turbo Quant - off

# Context: Code (Mixed)

# Single request results

Test TTFT(ms) TPOT(ms) ppTPS tgTPS E2E(s) Throughput PeakMem

pp 1024 / tg 512 1202.8 15.1 851.3 66.2 8.9 171.7 22.4 GB

pp 4096 / tg 512 6408.2 14.1 639.2 70.9 13.6 337.8 23.5 GB

pp 8192 / tg 512 14886.6 18.4 550.3 54.4 24.3 358.0 24.1 GB

pp 16384 / tg 512 28808.4 21.4 568.7 46.8 39.8 424.7 25.4 GB

pp 32768 / tg 512 56761.6 22.0 577.3 45.6 68.0 489.2 28.0 GB

pp 65536 / tg 512 131379.6 22.2 498.8 45.0 142.8 462.5 33.2 GB

pp 131072 / tg 512 341823.6 29.7 383.4 33.8 357.0 368.5 43.9 GB

pp 200000 / tg 512 656975.3 40.9 304.4 24.5 678.0 295.8 54.9 GB

# Batch results

Batch tgTPS ppTPS avgTTFT(ms) E2E(s) Speedup

1x baseline 66.2 851.3 1202.8 8.9 1.00x

2x 55.7 611.1 2670.3 21.7 0.84x

Quality testing:

- Seraphim Serapis Tool-Eval-Bench 89/100 - same as 8 bit version

- 0rand/DragonScale Bench - 98/100 - very similar quality as DeepSeek v4 Flash 0731 (4bit/8bit) on 2xDGX Spark Cluster and OpenAI GPT 5.6 Luna

reddit.com
u/0rand — 3 days ago