New to local AI world (M2 Max 64gb)
Hello im new to local ai and i just want to get the most information possible about it
For context im a software developper that never really used ai before and what to give it a try with my new machine (so basically im trying to achieve agentic result)
I already have omlx and some model installed
But im still trying to understanding all the concept and specification relatated to mlx etc
I understand that quantification shrink the model size trying to keep every essential data to run faster
Still dont know if their is really any advantage to not run every model in their 4bit variant ?
I learned that their is some model training on basic model idk if their are some reputated training like heratic, uncensored, abliterate or some name like opus etc
I see some like quantification method i guess like AXQ optiQ oQ etc idk what is it and the better one
I know the goal of A3B look like the best type of mode to get in every circomstence (maybe im wrong) with the most adequat part of model being active to answer the question)
And idk the best models im currently trying:
- gemma-4-26B-A4B-it-qat-OptiQ-4bit
- Qwen3.6-35B-A3B-OptiQ-4bit
- Qwen3-Coder-30B-A3B-Instruct-MLX-4bit
Maybe i will need some smaller model for lighter task i really dont know
I see some people here talk about fp16 for m1/m2 serie why ?
Any help is appreciate^^
Thanks for reading this