Anthropic's Model 2 story may be less interesting than its benchmark problem
Anthropic's latest risk report has been getting attention because it discusses an unreleased internal model called Model 2.
According to the report, Model 2 is slightly more capable overall than Mythos 5 and is being used internally for coding, research, data generation and agentic work. Anthropic currently doesn't plan to release it externally.
But I think there's a more interesting part of the report.
Anthropic says some of its most concrete task-based evaluations have saturated and are no longer adequately capturing increases in model capabilities.
That's a pretty fundamental problem.
A benchmark is useful when the score changes meaningfully as capability changes.
But imagine a benchmark where:
Model A = 80
Model B = 84
You conclude B is only slightly better.
Now imagine B has developed capabilities the benchmark wasn't designed to detect.
The 4-point difference might dramatically understate the actual difference.
Anthropic also says it has become less confident in some of its risk assessments for this reason, while seeing early signs of accelerated AI R&D.
So I'm wondering whether the next major bottleneck in frontier AI is actually evaluation.
Not:
“how do we build a smarter model?”
but:
“how do we discover capabilities that our existing tests weren't designed to see?”
I'd be interested in hearing from people working with evaluations, agents or red-teaming.
What kind of evaluation would you trust more than a static benchmark once models start saturating the benchmark itself?