
Game over. 22GB local models run in Pi now outperform Claude Code Opus 5 High on real-world coding tasks published after training cutoffs
Ran this benchmark on recently published real code base benchmarks to test the Sharp chat template that reduces token use and fixes bugs on locally run Qwen3.x models. Thought I’d bench Opus 5 high and Sonnet 5 medium alongside, for fun. I guess we have finally reached the point where the reduction in Claude’s quality has finally surpassed the upwards trend of local models for actual real world work. Claude Max 20x subscription btw. Not for long though hahah.
I don’t care about Artificial Analysis index or published benchmark numbers. If Opus is beaten by the models I run on my own computers when it comes to fixing real bugs without introducing regressions in real life code bases, it doesn’t matter if it’s because Anthropic is silently reducing Opus quality to sell more Fable tokens, or whatever is going on. EDIT: Someone asked me to add the chat template link to the op, so: https://huggingface.co/peculiar-ragdoll/Qwen-Sharp-Chat-Templates The "Sharp Qwen3.8-27B" model is here: https://huggingface.co/peculiar-ragdoll/Dirk-Qwen3.8-27B-GGUF and "Nail (Sharp 35B-A3B)" is here: https://huggingface.co/peculiar-ragdoll/Nail-Qwen3.6-35B-A3B-GGUF