


When will scicode be saturated?
SciCode
HLE
CritPt
It has become my favorite benchmark since it is a fair test of how models perform in science when allowed to use coding which is their strong side.
But it has been so slow.
As you can see, CritPt has progress in the shape of a box, but seriously their progress rate is difficult to measure so I just toon the progress from the moment the models started getting good to the current plateau (GPT 5.6)
HLE will be counted from January 2025 release of DeepSeek R1 to Claude Opus 5.
SciCode counted from Claude 2.0 to Claude Fable 5
improvement% per month
CritPt ~2.3%
HLE ~2.6%
SciCode ~1.2%