Image 1 — GLM-5.3 is out on AA, and I'm fed up with their Intelligence/cost plot
Image 2 — GLM-5.3 is out on AA, and I'm fed up with their Intelligence/cost plot
▲ 20 r/ZaiGLM+2 crossposts

GLM-5.3 is out on AA, and I'm fed up with their Intelligence/cost plot

I think AA's intelligence/cost plot is seriously misleading, so I decided to make my own. Their plot is in the second image.

All points are at max thinking. All intelligence index scores are from AA. All cost scores are from AA too except where noted below.

What changes between AA's plot and mine:

  • Changed X scale from logarithmic to linear, because people's money is not logarithmic
  • Added DeepSeek V4 Flash 0731 as it is priced today by third party providers on OpenRouter (note: you don't get this today with OpenCode Go/Zen, but it's been promised you will soon).
  • Added GLM-5.3 as it will be priced by third party providers on OpenRouter in <2 weeks, assuming no license changes from 5.2. Note: you don't get this on OpenCode Go/Zen.
  • Added Qwen3.8-27B. Cost per task was crudely calculated from
    • 47,166 output tok/task (AA)
    • tg 55 tok/s @ 350W, as crudely observed on my RTX3090 (IQ4_XS shows negligible quality loss - dedicated post coming soon)
    • today's US median residential electricity price
    • today's UK median residential electricity price + today's GBP/USD fx
    • +15% (finger-in-the-air) for prefill and waiting for tools
    • Hardware priced at 0, on the basis that both a RTX 3090 PC and a 64GB Strix Halo are desirable gaming/work machines anyways.
    • These maths are meant to produce a rough back-of-the-envelope figure and should not be taken authoritatively were you to zoom into the bottom-left corner of the chart. They don't want to answer how much cheaper it is to run Qwen at home vs. DSv4 on OpenRouter, because they are both so cheap that the difference is inconsequential for most of the population.

Note: not including the cost of hardware stops being defensible once you upgrade to a 128GB Strix Halo (almost nobody needs that much RAM if not for AI). This is why I did not add self-hosted DeepSeek IQ2_XXS to the chart; it would likely also sit lower on the intelligence axis than the MXFP4 native model. Same argument for a ~$16k rig needed to run GLM-5.3 IQ4 locally. I'm not saying they're not worth the expense (privacy is priceless), just that pegging them on the plot is a much more nuanced exercise.

u/crusaderky — 18 hours ago
▲ 111 r/huggingface+1 crossposts

LFM2.5-2.6B model+KV cache quantization report

LFM2.5-2.6B is a new tiny model by LiquidAI, with benchmarks that put it head to head with much larger models.

I've run llama-perplexity on many model GGUF quants, crossed with many KV cache quants, to understand the model's best overall quantization for any given amount of memory.

I also show how different quantization metrics show (or hide) model degradation.

Full report and commentary

Interactive HTML plots

If you don't have time to read

  • The model fits on an 8GB Raspberry Pi with no material degradation and on a 4GB Raspberry Pi with contained degradation.
  • DO NOT use Q4_K_M.
  • On this model, model quant quality degrades faster than KV cache quant.
  • Abliteration comes with a flat cost of ~0.075 KLD.
  • Logarithmic KLD and Top-1% plots lie to you by telling you that quality degradation is smooth, while it's actually a cliff.
u/crusaderky — 13 days ago

Nanbeige4.2-3B: I'm not impressed

I've tested Nanbeige-4.2-3B. On paper, the benchmarks promise it blows away Qwen3.5-9B and Gemma4-12B.

My goal was to have something very light and fast to replace Qwen3.6-35B (or finetunes thereof) for simple and straightforward coding tasks. In the past I tried downgrading Qwen3.5-9B and it was not good enough to be considered.

The model is currently broken in llamacpp master - this PR fixes it:
https://github.com/ggml-org/llama.cpp/pull/26324

After fixing its issues, I played around with it and must say I'm not impressed.

To begin with, it's a looped model: all layers are traversed twice. This means that, at a theoretical baseline, it has the speed and context size of a 6B model. It's nice to be able to run the weights at a Q6 quant and barely notice the size difference from Q4, but you will have to compensate by using a very bad KV cache quant, because the context is enormous for the size. 128k of kvarn3 t2048 context, which I must point out is both very tight and at the edge of the cliff of what is usable without extreme degradation, costs 5.2GB. That's ginormous for a model this size. 256k kvarn5 won't fit on 16GB VRAM after you factor in weights and desktop.

The model uses the same "hack" to get good benchmark results that Laguna-S-2.1 uses: at [max] thinking level, where it is benchmarked, it thinks and thinks and thinks and just does not stop. This means that, besides being atrociously slow (wall time per task), it burns through its context budget VERY fast even for simple tasks.

I gave it two very straightforward, uncomplicated brownfield maintenance tasks in a project with a robust AGENTS.md and skills. It flunked both.

The only good thing I have to say is that tool calling is rock solid. After the llamacpp PR above, it never fails a single tool call.

Is it actually better than Qwen3.5 9B? Hard to say: I've only had bad experiences with that too and I have a hard time telling apart models that consistently fail at the simplest tasks. Worth noting that Nanbeige has exactly the same size in memory (at 128k) and same speed.

time-per-task, Qwen3.6-35B-A3B with experts spilled to host memory is vastly faster and actually produces correct outputs.

Want something small? Not a good model (tiny on disk, enormous in VRAM).

Want something fast? Also no, particularly when you measure time-per-task instead of tok/s.

Want something precise and reliable for the very easy stuff? Also no.

u/crusaderky — 21 days ago

What's the carrying capacity of a mule?

A mule doesn't have a statblock. Given the price difference from a horse it would be sensible to say STR+4, speed 25~30. That would give it a bulk limit of (10+4) x2 for large creature = 28.

Two saddlebags have bulk L and carry 6 bulk, 2 of which does not get counted, like backpacks. Unlike backpacks, you should be able to load multiple pairs on the same animal as long as you're not riding them. IRL pic above for example.

So you should be able to load 7 pairs of saddlebags to carry 6*7=42 bulk, which will weight 4.1*7=28.7 -> 41 bulk total load (the last saddlebag can't be completely full; leave some spare room for barding).

I could not find anything about getting fatigue/exhaustion when travelling overland while encumbered (and any loaded pack mule is pretty much the textbook image that comes up when you think of "encumbered"). You do get a speed drop of course.

Total cost 2+0.2*7=3.4 GP, 41 bulk, speed at full load 15~20ft.

Splurge 6 GP extra for a horse to get +10ft speed.

A wagon with two horses will cost you 25+8*2 = 41 GP, carry 180~190 bulk, need a dedicated driver, and will be frequently problematic when going off the beaten path.

Did I get everything right?

u/crusaderky — 1 month ago

I mapped the KLD of KV cache quantization for Qwen3.6-35B-A3B and Gemma4-E2B QAT

TL;DR version

  • q8/q8 is nearly free on both models
  • q4/q4 is useable on Qwen and catastrophic on Gemma
  • turbo4 is sometimes slightly better, sometimes slightly worse, than q4_0
  • turbo3 and turbo2 allow compressing the cache to unprecedented levels - but you'll pay dearly for it
  • K is sometimes more sensitive than V, sometimes less, sometimes they're symmetrical

Full analysis

Nuance, caveats, zoomable plots, and the software to replicate these plots with any model:

https://github.com/crusaderky/pixi-llm-recipes/tree/main/perplexity#readme

u/crusaderky — 2 months ago