




My new benchmark. 3.6 Flash destroyed all of the other AI models.
Yeah granted my art skills aren't that great..
I used Gemini 3.6 Flash (no thinking) and compared it to Sonnet 5 (Medium), 5.6 Terra ( in the default chat mode, no thinking), Grok 4.5 (with default thinking). All of them are in the default configuration on each app, meaning no instruction, no thinking and just using the base model.
My prompt was the following: Imagine the 5 lines on this road are speed bumps and a car is about to cross it (holy english grammar sry). How many countable bumps will the driver feel?"
Obviously the trick is that the speed bumps are diagonal, and the correct answer should be between 10 (on the most ideal conditions, with horizontal speedbumps and the car going straight, with two wheel axles) and 20 (if the speed bumps or the vehicle's direction aren't horizontal/straight).
On the first shot, 3.6 Flash (not even 3.7 Flash, nor Pro nor any thinking mode) nailed it. It gave me the correct answer without any hesitation.
The 3 other AI models, Grok, Sonnet and GPT, all failed miserably, all of them didn't take into consideration the diagonal configuration of the speed bumps and the fact that there are 2 wheel axles.
I kind of take it as a great benchmark to test the ability of each model for their vision capabilities and link what they see with their reasoning. Gemini somehow clearly won that here. Full Gemini answer: https://share.gemini.google/ClKuRlRI2msZ