r/AMD_MI300

Optimizing Qwen3.8-27B on one MI300X with an open-source agent toolkit: 311 to 495 tok/s
▲ 46 r/AMD_MI300+3 crossposts

Optimizing Qwen3.8-27B on one MI300X with an open-source agent toolkit: 311 to 495 tok/s

Presets is an open-source toolkit for optimizing inference with agents, which we build at dstack. Here's one example of using it on a single MI300X.

Qwen3.8-27B went from 311 to 495 tok/s, +59%, at the full 1M context with p50 TTFT under 1.5s and four concurrent users at 10k in / 1.5k out.

The gains came from linked optimization sessions and source-level patches to SGLang's AITER attention backend.

What comes out is a portable preset that deploys on any AMD cloud, Kubernetes cluster, or bare-metal fleet: https://dstack.ai/blog/presets/

u/cheptsov — 24 hours ago