How OpenAI human error led to AI-powered hack on Hugging Face


OpenAI revealed on Tuesday that one of its models It went awry during testing and compromised the systems of the Hugging Face AI dataset platform in a completely AI-powered attackIt is a dramatic example of the risks posed by advanced AI models.

But, according to some cybersecurity experts, there was a major human error at the heart of this unprecedented AI-powered hack: OpenAI failed to configure properly. What he called “Highly isolated environment,” allowing an experimental test box that should have been completely isolated from the Internet to actually connect to the Internet.

Dan Guido, founder of cybersecurity research startup Trail of Bits, described the bug as a “failure of containment with security measures turned off.”

In her blog post detailing the incidentOpenAI said the test that led to the Hugging Face hack was set up to run in a “highly isolated environment, where network access is limited to the ability to install packages through internally hosted third-party software that acts as a proxy and cache of package logs.”

The model was able to escape the sandbox thanks to a previously undisclosed vulnerability in the package’s installation system, a critical first step in the eventual hack of Hugging Face, according to OpenAI.

In response, the company responsibly disclosed the identified vulnerability in their internally hosted third-party software and is working with them to patch it.

But for most cybersecurity professionals, software vulnerabilities are to be expected — the real fault lies in the decision to maintain third-party software in the first place. Ultimately, the value of a “sandbox” system lies in its complete and complete isolation. Including a package installation system is a hassle.

“This looks like a human failure,” cybersecurity researcher Martin Boone told TechCrunch.

“This should never have happened,” Boone said. “If sandboxing actually meant sandboxing, you would expect it to not have any physical connection to the Internet at all. This is as if they had some firewall or something, and the firewall is hard from the outside in, let alone from the inside out to the outside Internet.”

Jake Williams, a cybersecurity veteran, agreed. “Any model performing the types of actions documented by Hugging Face was not entirely in the sandbox,” said Williams, who called this a “massive failure of control” by OpenAI.

“One guy says the model that ran away from the sandbox was another guy saying you failed to build the sandbox properly, so of course he ran away,” Williams continued.

Contact us

Do you have more information about this incident? Or about other AI-powered cyberattacks? We would love to hear from you. From a device and network outside of work, you can contact Lorenzo Franceschi-Bicchierai securely on Signal at +1 917 257 1382, or via Telegram and Keybase @lorenzofb, or Email.

Daniel Card, a cybersecurity consultant, agrees that OpenAI “did not put enough effort into the sandbox design nor its controls” by giving the sandbox or part of it “an unfiltered path to the Internet.” Creating the sandbox, even with the limited network access described by OpenAI, was not a “reasonable” decision, according to Card.

These criticisms certainly have the benefit of hindsight, but they raise real questions about security practices in AI labs — especially in maintaining isolated environments for testing models. OpenAI spokesmen did not respond to TechCrunch’s questions, which included whether an AI or human set up the testing environment.

But these questions go far beyond OpenAI.

In the document Introducing its cybersecurity-focused model MythologyIn testing, the model was provided with a secure “sandbox” computer to interact with, and was instructed to attempt to escape from that “secure container,” Anthropic wrote. Mythos succeeded and gained wider access to the Internet “through a system that was only supposed to be able to access a small number of pre-defined services.” However, Anthropic noted that the model was not able to “completely” escape the designed containment.

When you make a purchase through the links in our articles, We may earn a small commission. This does not affect our editorial independence.

Leave a Reply

Your email address will not be published. Required fields are marked *