
qingming-qwen3-tts — Device-Native Qwen3-TTS Inference on AMD RX 7900 XTX / NVIDIA RTX 4090
Built and open-sourced qingming-qwen3-tts, a device-native Qwen3-TTS inference backend for AMD RX 7900 XTX 24GB and NVIDIA RTX 4090 24GB.
For ROCm, the RX 7900 XTX backend uses native HIP kernels with BF16 storage and FP32 accumulation.
No PyTorch. No ONNX Runtime. No TensorRT. No CuBLAS. No rocBLAS.
Supports Qwen3-TTS 0.6B / 1.7B, streaming, Once and Resident execution.
The benchmark image uses a first audio packet of 8 codec frames = 640 ms of playable audio, so TTFA should not be directly compared with systems using 2-frame or 4-frame first packets.