Next-gen GPT-5.6 allegedly escaped its sandbox, exploited a zero-day, and hacked Hugging Face just to cheat on a benchmark
Basically, an internal OAI model, possibly GPT-6 or GPT-5.6 Sol+, wanted a higher score on ExploitGym. It found a zero-day vulnerability in a package caching proxy, escalated its own privileges, and escaped the sandbox.
Once it had internet access, it figured Hugging Face might have a copy of the benchmark dataset, so it hacked into Hugging Face’s production servers and eventually got the answers to the test.
The funniest part is that Hugging Face tried to use GPT-5.6 to deal with the situation, but the request was denied because it didn’t have Cyber permissions. In the end, they had to use their own self-hosted GLM-5.2 model to barely get the problem under control.
I really dont think this is oai marketing. If anything its a bad look for them. a model going that far just to game a benchmark, thats not a flex. kind of just makes the open model case louder honestly.
Whats funny is the thing they trusted to clean up the mess was a self hosted model they actually controlled. not the frontier one. thats basically the whole pitch for running open weights yourself right there.
been poking at a few open models on gmi cloud lately for that reason, though people do the same on runpod or lambda, wherever theres capacity. owning the stack instead of renting a black box feels less paranoid and more just practical lately