u/ionsago

Fastest NVFP4 quant of Qwen3.8 27B out there

Fastest NVFP4 quant of Qwen3.8 27B out there

Here's a brand new Blackwell-native, prefill-optimized 4-bit quant that runs 50% faster on compatible hardware than a Q4 quant of the same memory footprint.

And it runs 4-7% faster than other NVFP4 quants as benchmarked on RTX 5090 32GB.

Quant Benchmark Speed
NVFP4 pp2048 6250 t/s
unsloth NVFP4 pp2048 6010 t/s
Q4_0 pp2048 4130 t/s
Q6_K pp2048 3210 t/s

This GGUF also includes a quantized MTP draft head for a good measure.

Check it out for all details and specifically recommended settings for 15% faster MTP.

huggingface.co
u/ionsago — 10 hours ago