Ran chinese models against my claude/gpt setup for a few weeks and the spend gap is wild
Api bills started getting stupid the past few months, like genuinely looking at my monthly spend and wondering if i am doing something wrong. Decided to run my own tests on chinese models instead of trusting whatever chart someone posts on twitter that week.
Deepseek, qwen, kimi went up against my normal claude/gemini/gpt rotation. Glm-5.3 got added this week when i finally got around to the new release so its early days for that one.
The spend gap is wild. Quality gap exists of course but its not anywhere near what the pricing makes it look, especially on iterative stuff where i am running the same task 5 times to get it right.
Closed models still win on hard reasoning most of the time. Once a prompt gets complicated with a bunch of conditions stacked deepseek and the older chinese ones start fumbling somewhere. Glm-5.3 actually held up better than i expected, felt closer to opus on a few of my tests but i will need more time before i say anything strong.
Claude and gpt still get my real work. Iteration heavy stuff just makes more sense on the lighter side because i am not burning premium tokens on a model to write the same function 4 different ways.
Would rather read other peoples actual usage notes than argue about charts at this point.