
How AI guardrails are impeding the work of offensive cybersecurity researchers

AI guardrails designed to prevent malicious use are now impeding legitimate offensive cybersecurity researchers, according to a TechCrunch report.
The article highlights how programs like Anthropic‘s Mythos and OpenAI‘s Trusted Access for Cyber, intended to vet users, are criticized for arbitrary restrictions.
Researchers such as Mark Dowd, Chris Anley, and Paolo Stagno argue that these guardrails hinder vulnerability discovery and exploit development, which are essential for defense.
Anley notes that the same tool can be both offensive and defensive, akin to a hammer.
Stagno avoids using frontier models for vulnerability work due to data leakage risks, relying instead on open-source models run locally.
Anonymous researcher at a smartphone-component manufacturer finds Anthropic‘s tools unusable for security work.
Chris Thompson of RemoteThreat observes inconsistent guardrails even within vetted programs, pushing researchers toward Chinese open-source models like GLM.
He calls for responsible access rather than tighter restrictions, warning that defenders are being stifled as attacks increase in speed and scale.
Some researchers, like Giuseppe Cali, do not use AI for offensive work and thus are not affected.
The article underscores a growing tension between AI safety measures and the needs of cybersecurity professionals.


