How AI guardrails hinder the work of offensive cybersecurity researchers


For months, AI giants have created special vetted software and strict guardrails to limit the use of their models by malicious hackers. But these limitations now hamper the work of legitimate network defenders, as well as the work of offensive cybersecurity researchers.

In June, the US government It imposed export control restrictions Based on popular AI models from Anthropic Mythos and Fable. The move was prompted at least in part by a report that claimed it was possible to bypass template firewalls designed to prevent users from using them to build and carry out malicious cyberattacks.

Regardless of whether the incident was genuinely motivated Fears of escaping from prisonThe truth is that anthropic It has been marketed repeatedly Myths as A kind of internet doomsday machine Which can only be given to carefully vetted users, and even then with strict guardrails. (Export controls on Fable 5 and Mythos 5 have since been lifted. Fable 5 returned to public access on July 1; Mythos 5 was only resubmitted to US organizations that were vetted as part of the government review process.)

This type of gatekeeping is not unique to Mythos. Anthropic, its other models, and OpenAI all offer programs for cybersecurity researchers that they can apply to be vetted and — if approved — gain access to models with fewer cybersecurity restrictions: OpenAI’s Reliable access to cyber And the anthropic Cyber ​​verification program.

These guardrails have been widely criticized, particularly by researchers whose job is to find unknown vulnerabilities in systems and devise ways to exploit them before criminals do.

During a recent appearance on a cybersecurity podcast, Mark Dodd, a well-known security researcher, said: He said That “it’s not really comfortable for me to have these big arbitrary companies making arbitrary decisions about what’s safe and what’s not safe.”

Dodd has spent decades Find “zero days” and sell them – Previously unknown software flaws and the exploits that exploit them – to Western governments, rather than being reported to software makers so they can be patched. Governments pay a premium for vulnerabilities precisely because they remain open, which is beneficial for intelligence operations.

Dodd admitted that his work may make him biased, but he is not alone. Several people who work in offensive cybersecurity — which proactively scans systems for vulnerabilities — described to TechCrunch how they use AI tools and get around their own guardrails.

Chris Anley, chief scientist at security consulting giant NCC Group, said asking an AI model to try to exploit the flaw is a key step in ensuring it’s a real vulnerability worth fixing. But if the guardrail prompts the model to explicitly refuse to answer the question, the guardrail hurts advocates, he said.

“This is where the whole offensive versus defensive and guardrail part comes in, because ‘fix this code’ as a prompt is a core defense mechanism but also a roadmap for finding critical vulnerabilities in the code base,” Anley said. “So, at the same time, the same tool is an offensive tool and a defensive tool, and you can’t really get rid of both.”

“It’s like a hammer,” he continued. “You can’t build a house without a hammer. It’s certainly a tool, but it’s also an irreducible weapon.”

When he and his colleagues encounter such an obstacle, they sometimes turn to open-source AI models that have no guardrails at all.

Paolo Stagno, chief technology officer at CrowdFense, a well-known company that develops, acquires and sells unknown vulnerabilities to government agencies, agreed with Dodd, saying that AI companies “essentially treat customers like children who need babysitting” with their software and vetted guardrails.

Stagno said he and his colleagues use parametric models, but only for reverse engineering. He said they avoid using AI to help find vulnerabilities or create exploits, because feeding that work into a cloud-based model risks sensitive vulnerability data being leaked or absorbed into future training. He said that in this step they are using open source models that are run locally, as they do not depend on sharing data outside the model.

Giuseppe Cali, the security researcher who discovered zero days and developed the vulnerabilities, said guardrails do not hinder his work. This is because it does not use artificial intelligence in offensive work; Instead, he uses it for initial reverse engineering, to understand the code he’s analyzing, and to build support tools. For this reason, he said AI tools can speed up the process and allow him to focus on discovering vulnerabilities.

“I still want to own the actual bug detection and weaponization myself, and that won’t change if all the guardrails are lifted tomorrow,” Cali said. “I’m jealous of my mistakes, and I love this game too much to let the models play it for me.”

One researcher at a smartphone component manufacturer, who spoke on condition of anonymity because he was not authorized to speak to the press, said his employer is not part of Anthropic’s CVP program and, as a result, its tools are barely useful for finding vulnerabilities because the guardrails are so strict.

This person said: “If there is wind, we do anything related to security, but it stops and cannot be used.”

Chris Thompson — CEO of cybersecurity firm RemoteThreat and founder of Offensive AI Con, an event focused on offensive security and AI — said that in his experience using frontier AI models, guardrails can be inconsistent and operate differently every day. This is true even within the looser confines of the Anthropic and OpenAI programs examined.

“I think the practical effect is that you spend a lot of time negotiating with the model rather than working on the underlying security software,” Thompson said. “Instead of analyzing vulnerabilities and reasoning through exploitability, you’re trying to find why you’re getting inconsistent results or why models are over-sanitizing the output.”

As a result, researchers are relying on or being pushed toward open-source Chinese models, such as the GLM, which are freely downloadable and can be run locally without any scrutiny or restrictions on use, Thompson said.

“You have these responsible researchers who are being pushed away from U.S.-governed systems to foreign-owned systems,” he said. “I think putting these barriers in place is doing more harm than good.”

Instead of tightening restrictions further, Thompson called on frontier AI labs to open up their software, provide responsible access, as well as hold accountable those who misuse their tools. Otherwise, he said, advocates will lose the AI ​​race.

“There’s this big storm coming,” Thompson said. “There’s this big wave of attacks that are going to happen quickly and on an unprecedented scale.” “But the same security consulting firms and forensic researchers trying to make a difference are now being stifled.”

When you make a purchase through the links in our articles, We may earn a small commission. This does not affect our editorial independence.

Leave a Reply

Your email address will not be published. Required fields are marked *