Small model user here
I run small models awen3.5-9b and gemma-4-12b for some tasks. Recently tried GGUF format and looked pretty much okay. Getting 15-20 tok/sec and 100-150 on prefil with MTP.
Decide to try again oMLx. I cannot get better performance than ggufs. Maybe I am missing small models MTP versions. Or faster models. Or my settings bad?
So question: what models you guys using on similar hardware how much you are getting ?
Use cases: parsing, email drafting and general assistant.
My Device: mac m2 24gb ram