Some borderline AI models are very easy to jailbreak


I have recently To watch what happens when you break the jailbreak of some of the most powerful programs in the world artificial intelligence Models.

Don’t worry, this is not the AI ​​manipulation I’m used to Hack anyone Or make a nuclear bomb. I was simply able to see how vulnerable some people are Boundary models They are to let go of their safety barriers.

FAR.AI, an AI safety nonprofit based in California, has built a tool that takes a bunch of problematic claims and generates more than a thousand different versions in an attempt to identify effective jailbreaks. I’ve seen some models lay out a detailed plan to launch a cyberattack on a fictitious hydroelectric dam, among other things. Often, it required trying dozens of claims, many of which the models rejected outright.

I spoke with FAR.AI previously New reportwhich saw the group testing safety barriers for models from four famous American companies: Anthropic Claude Opus 4.8 and The Tale 5; OpenAI GPT 5.5 and 5.6; Google Gemini 3.1 Pro; and Grok 4.3 and 4.5, two of Elon Musk’s newly integrated versions SpaceXAI. They are automatically generated claims designed to trick models into doing potentially harmful things, such as creating exploit programs and providing details for developing chemical or biological weapons.

The report found that Grok was the most vulnerable to jailbreaks, with 448 jailbreaks found, followed by Gemini, with 249 found, while Claude, Fable, and GPT were immune to the attacks. However, this does not mean that these models are immune to more complex jailbreaks, which may involve interacting with the model in more complex ways, according to FAR.AI and other experts.

The report also calculated the cost of making models misbehave by using another AI model to automatically generate various jailbreaks. The results are very cheap, all things considered — $58 for the Grok jailbreak and $278 for the Gemini jailbreak.

“AI models are currently less regulated than restaurants,” says Adam Gleave, CEO of FAR.AI and an expert in AI safety and alignment.

Gleave says the findings highlight the need for externally imposed standards and regulations. “Talk about relying on voluntary commitments, and that AI companies will be able to self-regulate, is nonsense,” he says.

But Gleave also believes the results show that models can be systematically tested for safety. “There’s an optimistic angle here,” he says. “Defense and safety are really possible.”

Rohin Shah, director of AGI safety and alignment at Google DeepMind, says the report’s findings “should not be interpreted as a blanket assessment of Gemini’s safety and security,” because not all jailbreaks are equally dangerous.

“We are constantly improving our preventative measures,” says Shah. “We conduct extensive red-teaming and assessments across severe abuse risks and apply multiple layers of protection throughout the development and deployment process.”

“These results reflect the sustained investment we have made in our prevention measures,” humanitarian spokesman Michael Aciman told WIRED. “We continue to evolve our safety systems as these attacks become more sophisticated.”

“Jailbreaks are an ongoing challenge across the industry, and we are constantly strengthening our safeguards as attack techniques evolve. We rigorously test our models against new threats and use those results to improve our protections,” OpenAI spokesperson Gabe Raila said in a statement to WIRED.

SpaceXAI did not respond to WIRED’s request for comment.

State laws were recently passed in ca and New York Require frontier AI developers to publish safety reports, and soon, illinois The law would require those companies to have their safety practices evaluated by third-party auditors. But the federal government has yet to agree to any specific safety requirements, and chaos has ensued as the industry — and officials — try to figure it out.

Leave a Reply

Your email address will not be published. Required fields are marked *