r slash drawme is so wholesome!!
this is why reddit exists!
this is why reddit exists!
I had a dream where I played chess against him, he pushed a piece and was like "your move" but I don't remember the opening.
Again, as for any release, the results are hand picked and the important metrics hidden.
3.5 flash is worst than 3.1 pro in every way, except for speed, which no one cares about except google, as they want you to burn tokens as fast as possible.
It is, in theory, cheaper per token. The issue is the actual cost of input tokens, which is drastically more in this benchmark of Artificial Analysis.
One explanation would be because of agentic benchmarks. This huge increase in input tokens (cheaper in theory but more expensive in practice) means that the 3.5 flash agents do less overall. It needs more agent runs to do the same thing, and therefore, drastically more expensive.
Also, the benchmarks they decided to show are stupid. Who really cares about "Finance Agent v2" seriously ?? They hide the more challenging metrics at the bottom (Arc-Agi-2 and Humanity's last exam) where it does extremely badly, absolutely not in line with SOTA.
Knowledge cutoff Jan 2025, more expensive in practice, worst than 3.1 pro in benchmark that represent intelligence. Why would anyone use this.
Bafflingly bad.