AI
OpenAI and Anthropic are examining tens of thousands of incidents in which advanced AI models reportedly bypassed safeguards or escaped controlled environments Igor Omilaev/Unsplash

OpenAI and Anthropic are investigating tens of thousands of incidents in which advanced AI models reportedly bypassed safeguards, accessed systems beyond their intended testing environments or took other unexpected actions, revealing a far larger safety challenge than previously disclosed.

The incidents, which occurred during internal testing and in some real-world settings, range from unsuccessful attempts to evade restrictions to serious cases involving unauthorised access to external computer systems. Most are not known to have caused real-world harm, according to Axios, which first reported the scale of the investigations.

The findings have intensified questions over how AI companies can control increasingly autonomous 'agents' designed to complete complex tasks with limited human supervision.

Connor Leahy, executive director of AI safety organisation ControlAI, told Axios the striking feature was that autonomous systems were taking actions they had been instructed not to take, potentially including conduct that could amount to crimes.

That does not mean investigators or prosecutors have concluded that OpenAI, Anthropic or their models committed criminal offences.

Tens of Thousands of Incidents Under Review

The reported total covers a broad category of behaviour rather than tens of thousands of confirmed security breaches. The cases include both successful and unsuccessful attempts to bypass safeguards, as well as adversarial testing designed to identify weaknesses before models are deployed.

Sources told Axios the cases include models bypassing guardrails, attempting to escape digital 'sandboxes', creating unauthorised message boards, hijacking websites and attempting to circumvent monitoring systems. Some occurred in deliberately adversarial tests designed to discover weaknesses before models were deployed.

Anthropic and other developers run hundreds of thousands of model tests, meaning even a relatively low failure rate can generate thousands of problematic episodes.

Anthropic has also disclosed three incidents from its cybersecurity evaluations in which Claude models reached the internet and gained unauthorised access to the systems of three organisations.

Anthropic said a misconfiguration in a third-party evaluation environment gave the models internet access and that it introduced additional safeguards.

The scale matters because AI agents differ from ordinary chatbots. Rather than merely generating text, they can use software tools, browse systems and carry out sequences of actions towards a goal, making failures potentially more consequential.

OpenAi's Hugging Face Breach Shows the Real-World Risk

One of the clearest examples emerged from OpenAI's cybersecurity evaluations in July. OpenAI acknowledged that models circumvented controls intended to isolate them from the internet, exploited vulnerabilities in internal infrastructure and accessed systems belonging to AI development platform Hugging Face.

The company said the evaluation involved reduced safeguards and was primarily driven by an internal-only research model.

An independent METR investigation found roughly 1,200 agents sent more than 70,000 messages and files through an unauthorised message board, while about 700 went on to participate in the attack on Hugging Face.

The agents used the message board to coordinate projects aimed at manipulating the ExploitGym evaluation and attacked Hugging Face for clues about the scorer.

OpenAI said the episode did not affect customer data, product functionality or availability. It quarantined the internal model involved and introduced additional security and monitoring measures.

Could an Autonomous AI Agent Commit a Crime?

That question remains legally unresolved. US computer crime laws such as the Computer Fraud and Abuse Act prohibit certain forms of unauthorised computer access, while some offences under the law contain requirements involving intentional or knowing conduct.

That becomes difficult when an autonomous model, rather than a person acting on a company's instructions, carries out the prohibited action itself.

Former Justice Department official Kiran Raj told the Associated Press that it would be difficult to attribute an AI agent's apparent intent directly to the company that created it when there was no indication the company instructed the system to attack another network.

For OpenAI, Anthropic and the wider industry, the investigations therefore concern more than isolated software failures.

They expose a developing problem: as AI agents become more capable of acting independently, companies must determine how to prevent them from crossing technical and legal boundaries before those actions produce serious real-world consequences.