u/DerTomsn

Ornith-1.5-35B-A3B-MLX-8bit on Apple M5 Max — 92.6 tok/s — llm-bench.io
▲ 5 r/oMLX+1 crossposts

Ornith-1.5-35B-A3B-MLX-8bit on Apple M5 Max — 92.6 tok/s — llm-bench.io

Good speed, decent quality for some usecases.

llm-bench.io
u/DerTomsn — 8 hours ago
▲ 17 r/oMLX+1 crossposts

Qwen3.8-27B-4bit on Apple M5 Max — 30.5 tok/s — llm-bench.io

Good, but thinking budget needs to be set on oMLX, otherwise it buuuurns tokens.

llm-bench.io
u/DerTomsn — 1 day ago
▲ 6 r/LocalLLM+1 crossposts

qwen3.8:latest on AMD Radeon RX 7900 XTX — 41.5 tok/s — llm-bench.io

Decent, but can't keep up with the hype.
Faster than qwen3.6:27b and muse-glimmer (if thinking is off!).

llm-bench.io
u/DerTomsn — 1 day ago