Qwen 3.8 27B and Deepseek V4 Flash. Why are we building data centers?
I feel like these 2 models have shown that massive models that require hundreds of thousands of dollars worth of compute are unnecessary. Sure, training these models takes a good bit of hardware, but running them can be done at the fraction of the investment of the trillion parameter class models.
GLM 5.3 might also fall into the same "reasonable" category, however, for small companies rather than individuals.
I think a qwen 3.8 120b MoE model would also be a good release for business use.