Anthropic Wants Governments to Stop 'Catastrophic' AI Models Before They Are Deployed
The Claude developer's proposed framework would introduce catastrophic-risk testing, independent evaluations and revenue-linked penalties for frontier AI companies

Anthropic wants governments to gain the power to stop advanced AI models from being released if they pose a significant risk of catastrophic harm, warning that rapidly improving systems could outpace regulators before safeguards are in place.
The Claude developer unveiled its Advanced AI Framework in June, calling for tougher oversight of frontier systems as their capabilities accelerate, and on Thursday released a threat-intelligence report detailing attempts to misuse Claude.
Under the proposal, governments could prevent or deter dangerous deployments, while developers could face escalating civil penalties linked to their global annual revenue for violations.
The company stressed that such powers would require safeguards against government overreach and should apply only to the most advanced AI developers.
Governments Could Block Dangerous Deployments
Anthropic's proposal goes beyond requiring AI companies to disclose how their models are tested.
The company wants frontier developers to conduct catastrophic-risk testing, publish safety information, and submit their models to qualified independent evaluators.
If testing identifies a significant risk of catastrophic harm, Anthropic argues that governments should have legal authority to prevent or deter the model from being deployed.
The company says that authority would go beyond the powers available under current US law and proposals before Congress.
Anthropic also recommends escalating civil penalties tied to global annual revenue for repeated violations, giving regulators financial enforcement powers alongside the ability to intervene before deployment.
We're publishing our most detailed threat intelligence report to date.
— Anthropic (@AnthropicAI) September 10, 2026
It covers how people tried to misuse Claude—for cyberattacks, influence operations, surveillance, biology, and building weapons—and how we found and stopped them.
We disrupted every operation in the report,…
Four Areas of Catastrophic Risks Are Identified
Anthropic's framework identifies four categories of risk that could justify heightened scrutiny: biological threats, cyberattacks, loss of control, and automated AI research and development.
The company warns that increasingly capable models could make developing biological weapons easier or help attackers identify vulnerabilities in critical infrastructure.
Another concern involves systems becoming difficult for their developers to control as their capabilities increase.
Anthropic also points to AI systems increasingly automating AI research itself, potentially accelerating improvements and amplifying other risks.
The company said its own Claude Mythos Preview discovered thousands of high-severity software vulnerabilities, including flaws affecting major operating systems and browsers, as evidence of how quickly capabilities are progressing.
Rules Would Target Only Frontier Developers
Anthropic is not proposing that every AI company face the same regulatory regime.
Its framework would apply to models trained using more than 10²⁵ floating-point operations and developed by companies earning more than $500 million in AI-related revenue or spending more than $1 billion on AI research and development.
Those thresholds are intended to focus regulation on companies developing the most computationally powerful systems while limiting the burden on smaller businesses and less capable models.
Anthropic says developers covered by the rules should also maintain robust security programmes to protect model weights and training infrastructure from cyberattacks and theft.
Anthropic Says Industry Cannot Police Itself
The proposal is notable because it comes from one of the companies developing the frontier systems that would face greater government oversight.
Anthropic argues AI companies should not have sole responsibility for deciding whether their own models are safe enough for release.
The company is also proposing investments in biological surveillance, critical infrastructure protection, and systems capable of detecting or responding to AI operating outside developers' control.
Anthropic acknowledged that questions surrounding advanced AI regulation remain complex and its proposals are likely to face debate.
But its central argument is that waiting for catastrophic capabilities to emerge before establishing government authority could leave regulators responding too late.
As frontier systems become more capable, Anthropic wants governments to have the legal tools to intervene before a dangerous model reaches the public.
© Copyright IBTimes 2026. All rights reserved.

























