Which model did this—or which architecture made it possible?
For the last few years, we have evaluated AI systems primarily by asking which model produced a result.
I suspect that question is beginning to lose some of its importance.
As models gain tools, memory, retrieval, evaluators, feedback loops, specialized roles and stopping conditions, the decisive unit is no longer the model alone. It is the harness: the architecture that determines what the model sees, what it may do, how its output is tested, what is remembered and when another iteration is justified.
The model will still matter. Different models—and combinations of models—will reveal very different strengths. But the model may increasingly become one component inside a larger cognitive system.
A weaker model inside a well-designed architecture might sometimes outperform a stronger model operating in a poor one.
So when an AI system produces an unexpected discovery, solves a difficult problem or shows something resembling emergence, will the important question still be:
“Which model did this?”
Or will it become:
“In which architecture did this become possible?”
Where do you think the decisive capability will come from—the model, the harness, or the interaction between both?