The AI guardrails cybersecurity researchers encounter daily are no longer a minor inconvenience: they are actively redirecting legitimate offensive security work away from US-governed frontier models and toward Chinese open-source alternatives with zero restrictions.

That is the blunt conclusion from several people working in offensive security who spoke about how AI tools fit into their workflows, and whose frustrations have sharpened following a turbulent few weeks for Anthropic’s flagship models.

The Export-Control Flashpoint That Shook the Field

Anthropic released Fable 5 and Mythos 5 on 9 June 2026, describing them as the most capable models in the company’s history, with advanced reasoning and code-analysis capabilities. Cloud Security Alliance Labs noted that the export-control directive from Commerce Secretary Howard Lutnick followed just three days later, on 12 June 2026.

The trigger, according to TechPolicy.Press, was information passed to officials from an Amazon researcher via CEO Andy Jassy, alleging a jailbreak within Fable 5 that could allow bad actors to conduct cyberattacks. Lutnick’s letter to Anthropic CEO Dario Amodei prohibited access ‘by any foreign national, whether inside or outside the United States,’ including foreign national Anthropic employees.

Anthropic suspended both models globally at 5:21 p.m. ET on 12 June. As FifthRow reported, the company cited the technical impossibility of filtering users by nationality in real time across dozens of global cloud platforms as the reason for a universal shutdown rather than a targeted one.

The mechanism was a Bureau of Industry and Security ‘Is Informed’ letter under the Export Control Reform Act, requiring an individually validated export licence, rather than a finalised rule under the Export Administration Regulations, according to Digital Applied. Fable 5 returned to general access on 1 July; Mythos 5 has been reintroduced only to vetted US organisations as part of the government’s review process.

AI Guardrails Cybersecurity Researchers Find Most Frustrating

The export-control episode is an extreme case, but it sits within a broader picture of friction that offensive security professionals describe as a daily grind.

Chris Thompson, chief executive of cybersecurity firm RemoteThreat and founder of Offensive AI Con, said the guardrails inside even the vetted programmes are inconsistent. ‘I think the practical impact is you spend a lot of time negotiating with the model instead of working on the core security program,’ he said. ‘Instead of analyzing a vulnerability and reasoning through the exploitability, you’re trying to find why you’re getting inconsistent results or why are models over-sanitizing the output.’

The consequence, Thompson said, is that researchers get pushed toward Chinese open-source models like GLM, which can be run locally with no vetting or usage restrictions. ‘You have these responsible researchers that are being pushed away from U.S.-governed systems to foreign-owned systems,’ he said.

Chris Anley, chief scientist at NCC Group, framed the underlying problem precisely. ‘Fix this code as a prompt is both an essential mechanism for defense but also a roadmap for finding critical vulnerabilities in the code base,’ he said. ‘So at the same time, the same tool is both an offensive tool and a defensive tool, and the two can’t really be unpicked.’ When guardrails block that prompt entirely, defenders lose just as much as attackers gain nothing.

Paolo Stagno, chief technology officer at Crowdfense, said his colleagues use frontier models only for reverse engineering and run open-source models locally for vulnerability research, specifically to avoid feeding sensitive data into cloud-based training pipelines. The vetted programmes, he said, treat customers ‘like children who need babysitting.’

Not everyone is equally frustrated. Giuseppe Cali, a security researcher who develops exploits, said guardrails have not impeded him because he does not use AI for offensive work in the first place. ‘I still want to own the actual bug discovery and weaponization myself and that wouldn’t change if all guardrails were lifted tomorrow,’ he said. ‘I am jealous of my bugs, and I like this game too much to let models play it for me.’

What the Vetted Programmes Actually Offer

Both Anthropic and OpenAI have attempted to thread the needle with application-based access. Anthropic’s Cyber Verification Programme covers its Opus and Sonnet models, is free to apply for, and aims to return a review decision within two business days, according to the Anthropic Claude Help Centre. It is not currently available on Amazon Bedrock or Google Vertex AI, and organisations using Zero Data Retention cannot participate without a separate workspace.

OpenAI’s equivalent, the Trusted Access for Cyber programme, has expanded to include a fine-tuned model, GPT-5.4-Cyber, built explicitly to reduce friction for security work after earlier GPT models refused dual-use queries from cyber partners. OpenAI is also committing $10 million in API credits through its Cybersecurity Grant Program to support defensive work on open-source software and critical infrastructure.

The intent is genuine. The execution, according to researchers actually inside these programmes, remains patchy. Thompson’s call is straightforward: open the programmes further, enforce accountability for misuse, and stop treating the people trying to find vulnerabilities before criminals do as the primary threat. The alternative, he argued, is handing those researchers to systems that answer to no one in Washington at all.

Share.

Marcus Hale has been filing general news for the better part of fifteen years. He started at a regional evening paper, moved to a mid-sized digital outlet covering UK news, and spent three years as a general assignment reporter before going freelance. He has covered inquests, council elections, infrastructure announcements, and the kind of stories that sit on page five but matter on page one. He writes about public services, housing, local government, and the institutional stories that take six months to develop and thirty seconds to read. He prefers facts to angles and considers that unfashionable. Marcus lives in Bristol. He still reads the local paper and thinks that makes him an endangered species.

Leave A Reply