Physical Address
304 North Cardinal St.
Dorchester Center, MA 02124
Physical Address
304 North Cardinal St.
Dorchester Center, MA 02124

OpenAI says its AI models were accidentally hacked by open source AI platform Hugging Face during internal testing. in Tuesday’s postOpenAI wrote that GPT-5.6 Sol and a “more capable pre-release model” discovered security vulnerabilities in their test environment, allowing them to access the Internet and target Hugging Face.
On July 16, face hug Detecting a security incident Which it says was driven by an “autonomous AI agent system.” Hugging Face’s AI agents detected and stopped the hack, which OpenAI has now acknowledged occurred while evaluating the cybersecurity capabilities of its models. “All evidence suggests that the models were hyper-focused on finding a solution for ExploitGym,” OpenAI says, a benchmark system that measures whether AI models can turn vulnerabilities into exploits.
As part of the effort to complete the assessment, the AI models gained access to the Internet by exploiting a zero-day vulnerability in a sandbox environment. From here, OpenAI says its models “concluded that Hugging Face likely hosts models, datasets, and solutions for ExploitGym,” and then “researched and successfully found ways to access confidential information that could be used to cheat the assessment:”
In one example, the model linked multiple attack vectors together, including the use of stolen credentials and zero-day vulnerabilities to find a remote code execution path on Hugging Face servers.
But despite the seriousness of this incident, OpenAI appears to be using the “unprecedented” attack as an opportunity to make its own AI systems look good — especially as it competes with cybersecurity rivals, such as Anthropic myths and Gemini Flash 3.5 Cyber. The OpenAI blog post has a diagram showing how GPT-5.6 Sol improves at maintaining multi-step cyber operations, and also encourages enterprise customers to Sign up to access its “cyber” security model.
OpenAI adds that it is now working with Hugging Face to investigate the security incident, and will implement new controls in its search environment.