
Ling 3.0 Flash on Strix Halo
vLLM ROCm/HiP, 4 bit compressed-tensors (int4)
Not a fair comparison, but Qwen-122b on the most optimized format possible I have run (rocmFP4) does not touch Ling in speed.
https://x.com/ciruai/status/2085996633267777554?s=46
Tool call is broken in certain harnesses. It works well with pi-type harnesses (omp, feynman). Has anyone noticed this?