u/anon_mistborn

▲ 4 r/oMLX

Ornith seems to be better.

Ornith seems to be better.

TL;DR: On an M3 Ultra, Ornith-1.5-35B-A3B (4-bit MLX) decodes 4.6× faster than Qwen3.8-27B (8-bit MLX) and scores slightly higher on a small hard eval. It also beats Qwen3.8-27B with speculative decoding, while running autoregressive.

Setup

  • Mac Studio, M3 Ultra, 256 GB unified memory
  • mlx-lm 0.31.3 / mlx 0.32.1
  • Ornith-1.5-35B-A3B, MLX 4-bit
    • 18.2 GiB download
    • 20.2 GB peak
  • Qwen3.8-27B, MLX 8-bit
    • 27.5 GiB download
    • 29.7 GB peak
  • Machine was shared and had other load. Numbers are a floor, not a best case.

Throughput

mlx_lm.benchmark -p 512 -g 512 -n 3, identical invocation for both:

Model Decode tok/s Prefill tok/s Peak mem
Ornith-1.5-35B-A3B 4-bit 107.9 2162 20.2 GB
Qwen3.8-27B 8-bit 23.4 408 29.7 GB

Trial spread was 1.5% and 0.35%, respectively.

Prefill is the bigger story: 5.3×.

A 20K-token prompt took Ornith ~25s end to end versus ~125s for Qwen3.8-27B. If your workload re-reads long contexts, that dominates.

It also beats Qwen with speculative decoding

I spent a day trying to make Qwen3.8-27B fast before testing Ornith:

Qwen3.8-27B 8-bit config Decode tok/s
Autoregressive 23.4
MTPLX, native MTP head, depth 2 65.7 (3.01×)
DFlash2 block-diffusion drafter, block 5 79.3 (3.37×)
Ornith, plain autoregressive 107.9

Both speculative stacks work and are genuinely impressive. DFlash2's 3.37× on Apple Silicon is close to its published 3.43× on an H200.

Ornith just beats them without needing either, with no drafter and no third-party runtime, because mlx-lm already ships qwen3_5_moe.py.

Quality: 12 hard cases, thinking enabled

Scoring is mechanical. Code tasks are executed against hidden assertions and pass only on a full suite.

Task Ornith-35B-A3B Qwen3.8-27B
code_exec (4, execution-scored) 4/4 4/4
multihop (3, two facts ~20K apart) 3/3 3/3
logic (3) 2/3 2/3
tool_schema (2, nested JSON) 2/2 1/2
Total 11/12 10/12
Wall time for the set 166s 498s

One logic item was ambiguous. Two vals gave the same "wrong" one, so discount it: 11/11 vs 10/11.

Qwen's other miss was invalid JSON on a nested tool call. For agent use, that is the failure mode that actually breaks loops.

Caveats, and they are not small

  • Not precision-matched. 4-bit vs 8-bit. Some of the gap is quantisation; the rest is likely 3B active vs 27B dense. I have not run the 4-bit Qwen control.
  • n=12. An 11 vs 10 spread is one item.
  • Vendor benchmarks disagree with me. On SWE-bench Pro, the only benchmark both publish, Qwen3.8-27B is ahead: 61.7 vs 59.6.
  • Thinking must be on. With enable_thinking: false, Ornith went 0/5 on arithmetic and recovered to 4/4 with it on. My first eval drew a conclusion that was purely an artifact of my own test design.
  • 122B comparison still running.

The bit that surprised me

MoE is not a handicap here. It is the reason this works.

With ~3B active parameters per token, memory tracks total parameters while speed tracks active parameters.

Ornith gets:

  • 4.6× the decode throughput
  • 5.3× the prefill throughput
  • 32% less peak memory

Also, Ornith-1.5 is architecturally Qwen's exact vocab size, i.e. a self-improvement-trained fork of Qwen's older MoE architecture.

Beating Qwen's newer dense model with it is a nice result for the training approach.

MIT licence, and it is multimodal.

reddit.com
u/anon_mistborn — 1 day ago