For months, the AI giants have developed specially vetted software and strict firewalls to limit the use of their models by malicious hackers. But these restrictions now hamper the work of legitimate network defenders as well as offensive cyber security researchers.
In June, the US Govt punched export control restrictions About Anthropic’s very popular AI models Mythos and Fable. The move was prompted at least in part by a report that claimed it was possible to bypass the security bars of models designed to prevent users from using them to build and carry out malicious cyberattacks.
Regardless of whether or not the event was truly motivating fear of jailbreakthe fact is that anthropic marketed many times Like a myth some sort of doomsday cyber machine it can only be given to carefully vetted users, and even then, strict safeguards are in place. (Export controls on Fable 5 and Mythos 5 have since been lifted. Fable 5 returned to general access on July 1; Mythos 5 was re-released only to vetted US organizations as part of the government’s review process.)
This kind of gatekeeping is not unique to Miphos. Both Anthropic and its other models, as well as OpenAI, offer programs that cybersecurity researchers can apply for verification and, if approved, access models with fewer cybersecurity restrictions: OpenAI’s Secure login for cyberware and anthropic Cyber Authentication Software.
These security bars have been widely criticized, especially by researchers whose job it is to find unknown vulnerabilities in systems and develop ways to exploit them before criminals do.
In a recent appearance on a cybersecurity podcast, noted security researcher Mark Dowd, he said that “I’m not very comfortable with these random big companies making arbitrary decisions about what’s safe and what’s not in security.”
Dowd spent decades finding and selling “zero days”. — previously unknown software flaws and exploits that exploit them — to Western governments instead of informing software makers for patches. Governments pay a premium for vulnerabilities because they remain open, which is beneficial for intelligence operations.
Dowd admits her work may make her biased, but she’s not alone. Several people working in the offensive cybersecurity field — who actively probe systems for vulnerabilities — described to TechCrunch how they use AI tools and deal with their defenders.
Asking an AI model to try to exploit a bug is a key step in confirming that it’s a real vulnerability worth fixing, said Chris Anley, chief scientist at security consulting giant NCC Group. But if the guardrail model prompts them to refuse to answer the question directly, it hurts defenders.
“This is where the whole attack defense and security part comes in, because ‘fix this code’ is both an important mechanism for defense and a roadmap for finding critical vulnerabilities in the code base,” Anley said. “So at the same time, the same tool is both an offensive tool and a defensive tool, and you can’t really choose between the two.”
“It’s like a hammer,” he said. “You can’t build a house without a hammer. It’s certainly a tool, but it’s also an inexhaustible weapon.”
When he and his colleagues face such a hurdle, they rely on open-source AI models that sometimes have no safeguards.
Paolo Stagno, chief technology officer of Crowdfense, a well-known company that develops, acquires and sells unknown vulnerabilities to government agencies, agreed with Dowd, saying that AI companies “basically treat customers like children who need babysitting” with their vetted programs and guardrails.
Stagno said he and his colleagues use boundary models — but only for reverse engineering. They’re reluctant to use artificial intelligence to help find vulnerabilities or create exploits, he said, because handing that work over to a cloud-based model risks leaking sensitive vulnerability data or being included in future training. For this step, he said, they use open-source models that run natively because they don’t rely on sharing data outside of the model.
Giuseppe Cali, a security researcher who found the zero days and developed the exploits, said it did not interfere with the work of the protectors. That’s because it doesn’t use AI for offensive work; Instead, it uses the initial reverse engineering, understanding the code it analyzes, and building supporting tools. For this, he said, artificial intelligence tools can speed up the process and focus on detecting vulnerabilities.
“I still want to own the actual bug detection and weaponization myself, and that won’t change if all the guardrails are lifted tomorrow,” Cali said. “I’m jealous of my mistakes and I love this game too much to let models play for me.”
A researcher at a smartphone component maker, who spoke on condition of anonymity because he was not authorized to speak to the press, said his employer is not part of Anthropic’s CVP program, and as a result its tools are barely useful for finding vulnerabilities because the safeguards are so tight.
“If the wind picks up, whatever we do with safety, it just stops and is unusable,” the person said.
Chris Thompson, CEO of cybersecurity firm RemoteThreat and founder of Offensive AI Con, an event focused on offensive security and artificial intelligence, said defenders in their experience using borderline AI models can be inconsistent and work differently from day to day. This is true even within the looser bounds of Anthropic and OpenAI’s vetted software.
“I think the practical effect is that you spend a lot of time negotiating the model instead of working on the underlying security software,” Thompson said. “Instead of analyzing the vulnerability and justifying it with exploitability, you’re trying to figure out why you’re getting inconsistent results or why the models are over-refining the product.”
As a result, researchers are relying on or pushing toward open-source Chinese models like GLM — freely downloadable models that can be run locally with no validation or usage restrictions — Thompson said.
“You have these responsible researchers being pushed from US-run systems to foreign-owned systems,” he said. “I think having these safeguards in place does more harm than good.”
Instead of tightening restrictions, Thompson urged AI frontier labs to open up their programs, provide responsible access and hold accountable those who misuse their tools. Otherwise, advocates will lose the AI race, he argued.
“There’s this big storm coming. There’s this big wave of attacks that’s going to happen at a speed and scale that’s never been seen before,” Thompson said. “But the same security consulting companies and legitimate researchers trying to make a difference are now being suffocated.”
When you purchase through links in our articles, we may earn a small commission. This does not affect our editorial independence.





