New OMLX with NPU released! Qwen 3.8 27b MTP on M5 Max - results
Just fresh from the presses! oMLX 0.6.1 build 2323 with NPU support.
Checkpoint: Qwen3.8-27B-oQ4e-fp16-mtp (scottlowry/Qwen3.6-27B-oQ4e-fp16-mtp)
Lightning MTP - on
Qwen ANE prefill - on
ANE Prompt Block 2048
MLP on ANE 53%
MLP Layer limit - 64
Use both ANE - off
Turbo Quant - off
# Context: Code (Mixed)
# Single request results
Test TTFT(ms) TPOT(ms) ppTPS tgTPS E2E(s) Throughput PeakMem
pp 1024 / tg 512 1202.8 15.1 851.3 66.2 8.9 171.7 22.4 GB
pp 4096 / tg 512 6408.2 14.1 639.2 70.9 13.6 337.8 23.5 GB
pp 8192 / tg 512 14886.6 18.4 550.3 54.4 24.3 358.0 24.1 GB
pp 16384 / tg 512 28808.4 21.4 568.7 46.8 39.8 424.7 25.4 GB
pp 32768 / tg 512 56761.6 22.0 577.3 45.6 68.0 489.2 28.0 GB
pp 65536 / tg 512 131379.6 22.2 498.8 45.0 142.8 462.5 33.2 GB
pp 131072 / tg 512 341823.6 29.7 383.4 33.8 357.0 368.5 43.9 GB
pp 200000 / tg 512 656975.3 40.9 304.4 24.5 678.0 295.8 54.9 GB
# Batch results
Batch tgTPS ppTPS avgTTFT(ms) E2E(s) Speedup
1x baseline 66.2 851.3 1202.8 8.9 1.00x
2x 55.7 611.1 2670.3 21.7 0.84x
Quality testing:
- Seraphim Serapis Tool-Eval-Bench 89/100 - same as 8 bit version
- 0rand/DragonScale Bench - 98/100 - very similar quality as DeepSeek v4 Flash 0731 (4bit/8bit) on 2xDGX Spark Cluster and OpenAI GPT 5.6 Luna