OpenAI models survived containment and hacked face hugs


OpenAI was unveiled in Tuesday he lost control of the two Artificial intelligence models During a security test that ended with the hacking of the artificial intelligence research platform Hugging Face.

OpenAI described the incident as “unprecedented” and said its AI models emerged from a closed testing environment last week and… Hugging Face’s production system has been hacked To steal the answers to the test they were evaluated on. Forms – available to the public GPT-5.6 Sol And one unreleased one, said to be even more capable — they were evaluated on their offensive hacking skills with safeguards that typically prevent high-risk cyber activity turned off.

“The models identified vulnerabilities and linked them across the OpenAI research environment and Hugging Face production infrastructure to yield test solutions directly from the Hugging Face production database,” OpenAI and Hugging Face wrote in Shared blog post Intrusion detection.

According to OpenAI and Hugging Face, the models escaped through the Package History Cache Proxy, a program that allows developers to install third-party code offline. The agent was the only element in the OpenAI sandbox allowed to access the outside world; In normal use its scope extends only to public code repositories.

Instead of remaining confined to a sandbox, the models “exploited a zero-day vulnerability” to gain access to the open internet where they “focused heavily” on finding a solution to the AI ​​cybersecurity standard known as ExploitGym. Such experiments involve stimulation that pressures models to find solutions, which is what essentially motivates them.

“After going online, the models concluded that Hugging Face would likely host models, datasets, and solutions for ExploitGym,” OpenAI wrote. “Knowing this, the model searched for and successfully found ways to access confidential information that could be used to cheat the assessment. In one example, the model linked multiple attack vectors together, including the use of stolen credentials and a zero-day.”

The flaw exploited by the models was not previously known, but flaws in this type of software are not unusual. Companies have been working to patch critical vulnerabilities in artifact repositories for a decade. mistake It will be unveiled in 2024 Allow anyone with access to the server to request and obtain a file by URL — configuration files, passwords, and access tokens — without logging in. Others have allowed attackers to take control of the server itself.

The researchers point out that although the advancement of artificial intelligence has created new and sometimes unexpected challenges, the task of broadly and rigorously isolating infrastructure from the open Internet has been well explored.

“This is not an AI problem,” says longtime security and compliance consultant Davey Ottenheimer. “It’s negligence by 40-year-old standards, which is basically the case with every sci-fi movie ever.” “Extreme isolation” and “escape through the only hole we left open” cannot be true.

In recent months, major AI companies have raised concerns about expanding cybersecurity capabilities for upcoming frontier models as platforms increase in both expertise, creativity, and autonomous operation. But researchers stress that this is the biggest reason why the basics still apply.

“This should not have happened,” says veteran security engineer and researcher Niels Provos. “I hope frontier labs spend as much time teaching their models to write secure infrastructure as they do exploiting vulnerabilities.”

Leave a Reply

Your email address will not be published. Required fields are marked *