u/Ok-Simple459

Qwen3.6 vs 3.5 on DGX Spark: identical throughput, except with one flag flipped
▲ 16 r/Qwen_AI+1 crossposts

Qwen3.6 vs 3.5 on DGX Spark: identical throughput, except with one flag flipped

Same config, same benchmark: Qwen3.6 performs within ±1% of Qwen3.5 across all four scenarios. Enable MTP and the 16-concurrent stress test jumps +24% throughput with −57% TTFT — exactly where speculative decoding is supposed to lose. Hypothesis: memory-bandwidth-bound serving on GB10 unified memory leaves compute headroom that MTP verification fills. 72.5% global acceptance rate, full breakdown by workload type.
https://docai.hu/en/blog/qwen36-mtp-gb10

u/Ok-Simple459 — 1 day ago
▲ 9 r/Vllm

Same config, same benchmark: Qwen3.6 performs within ±1% of Qwen3.5 across all four scenarios. Enable MTP and the 16-concurrent stress test jumps +24% throughput with −57% TTFT — exactly where speculative decoding is supposed to lose. Hypothesis: memory-bandwidth-bound serving on GB10 unified memory leaves compute headroom that MTP verification fills. 72.5% global acceptance rate, full breakdown by workload type.
https://docai.hu/en/blog/qwen36-mtp-gb10

reddit.com
u/Ok-Simple459 — 3 months ago