Opus 5, My experience so far
TLDR: Opus 5 High won against 5.6 Sol Ultra in my tests in terms of capability but my ChatGPT $100 plan still provides more value per dollar.
My test codebase is an earlier development fork of Axiom and it consists of 4 languages and 2 build systems: C, C++, Typescript and Python, CMake + internal TS transformer.
I have a list of known bugs from this fork that I have fixed in the current development branch and I had them documented by severity. The report itself is NOT present in the fork source tree or the machine I ran the tests on so agents couldn't have accessed it (unless they hacked my storage VPS protected behind SSH Keys + WireGuard + TOTP + Rosenpass).
I ran both Opus 5 High and 5.6 Sol Ultra on this fork with the same prompt "Scan the codebase and find potential interface bugs across the following boundaries: ...".
Opus 5 ran inside Claude Code for VSCode extension (menu didn't show an opus 5 option yet so I switch to it using command /model claude-opus-5, which Claude confirmed it switched to.) 5.6 Sol in ultra level ran inside Codex extension for VSCode.
Opus 5 found all of the bugs I had documented and a few more stability concerns. Which is highly impressive given these bugs took me two weeks of semi-automated and full manual testing to find out first.
5.6 Sol Ultra found a majority of the bugs but missed a few important but subtle ones. Claude certainly won here.
However Opus 5 already cranked up it's context past 90% by the time it was quarter done and was burning through tokens and consumed my 5-hour pro limit in 30 minutes and I had to use ~$30 worth of credits for the complete bug report. While 5.6 Sol Ultra only consumed 8% of my weekly max usage. (Note, to re emphasize, Claude was on pro $20 subscription and consumed the 5 HOUR limit in full + $30 credits, and Sol was on $100 ChatGPT subscription and consumed 8% of WEEKLY allowence).
I understand that comparing Pro and Max plans across vendors are not reliable but assuming a rough cost linearity, I can say even in the worst case Sol is measurably more token efficient than Opus.