OpenAI says Hugging Face was hacked through its pre-release models


OpenAI admitted on Tuesday that one of its AI models hacked into Hugging Face’s systems during an internal cybersecurity test that went wrong. Hugging the face at first The violation was returned To an “external AI agent.”

in A blog post published Tuesday afternoonOpenAI detailed the steps that led to models hacking the service.

“After investigation, we now know that this particular incident was driven by a combination of OpenAI models — including GPT-5.6 Sol and a more capable pre-release model, all while minimizing cyber denials for evaluation purposes — while being tested internally on a benchmark⁠ for cyber capabilities,” the post read.

In particular, the hack seems to have been focused on ExploitGymwhich is a publicly hosted benchmark for the ability of models to perform attacks based on found vulnerabilities. Standards like ExploitGym are commonly used in modular training to improve specific skills, but this is the first known incident where such testing has led to an actual cyberattack.

In this case, the model in question should not have access to the Internet, except for a specific tool that enables models to install the software packages they may need to complete their task. Instead, the model was able to find an undisclosed vulnerability in the package installer, which it used to access the wider Internet at will.

“Models was very focused on finding a solution for ExploitGym, doing their best to achieve a fairly narrow testing goal,” the OpenAI post said. “After accessing the Internet, the models concluded that Hugging Face likely hosted models, datasets, and solutions for ExploitGym. Knowing this, the models searched for and successfully found ways to access confidential information that could be used to cheat the assessment.”

Eventually, the models found vulnerabilities in Hugging Face’s infrastructure that allowed them to “obtain test solutions directly from the Hugging Face production database,” effectively providing answers to the benchmark.

For Hugging Face, the obvious result was a sophisticated and aggressive cyberattack, with “several thousand individual actions across a swarm of short-lived sandboxes, with autonomous command and control relayed over public services,” the company said in its initial disclosure.

OpenAI has identified and reported the vulnerabilities in the package installer, and is working with Hugging Face to further investigate the incident. The company also said it will implement new controls on both prototype testing and related infrastructure, with the aim of preventing similar incidents in the future.

It is unclear whether OpenAI will face any legal consequences as a result of the breach, although it is likely that the models’ actions may have violated the Computer Fraud and Abuse Act.

However, the result is an unusually vivid illustration of the power and dangers of frontier AI models operating over long time horizons. As OpenAI researcher Mika Carroll said Posted in response to the news“If this doesn’t convince you that incompatibility risks will be a major concern going forward, I don’t know what will.”

When you make a purchase through the links in our articles, We may earn a small commission. This does not affect our editorial independence.

Leave a Reply

Your email address will not be published. Required fields are marked *