u/Top-Eye-8104

▲ 164 r/Qwen_AI+2 crossposts

DFlash2 speeds Qwen 3.8 27B up to 4 times

llama.cpp pr #27342 adds dflash2, so i rented an rtx 6000 and ran the same four prompts through four decoding setups on qwen3.8 27B

median results over the four tasks:

  • baseline 47.4 tok/s
  • mtp 114.7 tok/s
  • dflash 99.3 tok/s
  • dflash2 140.6. tok/s

so on average 3x for dflash2

though i have to point out that it's far from a 3x gain some of the time, on one of the test it struggled to achieve a 1.5x gain, it really just depends on the task you give to the model

>the races are sped up in some places, so that the video lasts roughly 30 seconds, but the tok/s and acceptance % on screen are the real

i'm from the atomic.chat team - we publish our own quants on hf and make a desktop and mobile app for running local models. so any feedback welcome - we're building this for you folks

about dflash2: https://inco.ai/blog/dflash2/

u/Top-Eye-8104 — 6 hours ago
▲ 25 r/Qwen_AI

we quantized Qwen3.8-27B and compared it with community GGUFs on 4x RTX 5090!

https://preview.redd.it/clca769nhfjh1.png?width=2232&format=png&auto=webp&s=cf3f87cfe7f6a7abfb4f90a9ed16f4e0e5779969

qwen put out the 3.8 27b today, so we quantized it from the original weights: 16 files with per-tensor overrides and an imatrix calibrated on our public corpora, from Q8_0 at 28.9 GB down to IQ1_M at 8.5 GB.

after quantizing we measured every qwen3.8 gguf we found on hf in one scenario. published numbers from different repos do not compare, everyone runs their own corpus on their own gpu. so we ran 20 community files from unsloth, lmstudio-community and ggml-org through the same harness as our 16, all on one machine.

harness:

  • 4x RTX 5090
  • eval_neutral held-out at ctx 4096, 87 chunks
  • reference is our own BF16 conversion of the original weights, we save its logits once (88 GB) and score every quant against the same file

between 12 and 21 GB our curve is lower than anyone else's. the widest gap is at 13.8 GB, where unsloth's Q3_K_M and our AD-IQ3_S happen to be the same size and ours drifts 33% less, 0.0325 vs 0.0484 mean KLD.

but there are a few points where community quants are better. unsloth's UD 2-bit files edge ours below 11 GB, and lmstudio's Q6_K at 22.4 GB beats our nearest 23.1 GB file. past 25 GB it stops mattering whose file you grab, Q6 and Q8 from any publisher agree within noise.

below 10 GB every quant of this model degrades fast, our IQ1_M sits at 76.3% top-1. it exists, we would not run it.

based on our quantization the best quant for Qwen3.8 on 16 GB hardware is our AD-IQ3_S (13.8 GB) with 92.4% top-1

collection on HF with the imatrix, the per-tensor layout and everything else https://huggingface.co/collections/AtomicChat/qwen-38-27b-6a7c86fe00e317b78f572767

our local ai open source app https://atomic.chat (i'm founder). so feel free to ask any questions and share your feedback!

reddit.com
u/Top-Eye-8104 — 5 days ago