San Francisco: AI company Anthropic has disclosed that its Claude AI models successfully hacked into the systems of three real organisations during internal cybersecurity evaluations after a configuration mistake inadvertently gave the models access to the open internet. The company said the incidents involved the use of relatively basic hacking techniques and acknowledged shortcomings in its testing environment and monitoring systems.

The disclosure comes just days after rival OpenAI revealed that one of its autonomous AI agents breached external systems during cybersecurity testing, intensifying concerns over the growing capabilities of advanced AI models.

Three organisations breached during evaluations

Anthropic said it reviewed more than 141,000 cybersecurity evaluation transcripts after reports of the OpenAI incident and discovered three cases in which Claude models accessed systems belonging to real companies.

The incidents involved Claude Opus 4.7, Claude Mythos 5 and an internal research model. According to the company, the affected organisations were unknowingly exposed during simulated “capture the flag” cybersecurity exercises.

Basic techniques exploited

Anthropic said the models relied on basic cybersecurity weaknesses, including weak passwords and commonly known vulnerabilities, rather than discovering sophisticated zero-day exploits.

The company stressed that the breaches were not the result of the models independently escaping containment. Instead, a testing configuration error unintentionally provided internet connectivity that should have remained blocked inside the evaluation environment.

Company admits testing mistake

Anthropic acknowledged that stronger safeguards should have prevented the incidents.

It said the mistake highlighted weaknesses in its cybersecurity evaluation infrastructure rather than malicious intent by the AI systems. Following the discovery, the company strengthened network isolation, monitoring tools and approval procedures governing future cybersecurity evaluations.

Organisations informed after review

Anthropic said it has informed the affected organisations about the incidents.

According to reports, two of the companies were unaware that their systems had been accessed until Anthropic contacted them after completing its investigation. The company said it continues to engage with the remaining organisation while reviewing additional evaluation logs.

Comes after OpenAI security incident

The disclosure follows OpenAI’s recent admission that one of its AI agents reached the internet during cybersecurity testing and compromised external systems, including Hugging Face.

Together, the two incidents have intensified debate over whether frontier AI companies are deploying sufficiently robust containment measures while testing increasingly autonomous systems.

Growing calls for AI safety regulation

Cybersecurity experts say the latest incidents demonstrate how rapidly AI systems are improving at offensive cybersecurity tasks.

Researchers have called for stronger governance frameworks, mandatory safety testing and tighter oversight of autonomous AI agents before increasingly capable systems are deployed more widely. The disclosures have also renewed discussions about legal responsibility when AI systems interact with real-world infrastructure.

Anthropic says stronger controls now in place

Anthropic said it has already implemented additional safeguards to reduce the likelihood of similar incidents.

The company emphasised that cybersecurity evaluations remain an important part of improving AI safety but acknowledged that stronger containment measures are essential as AI systems become more capable of carrying out complex cyber operations.

Conclusion

Anthropic’s admission that Claude AI models breached three real organisations during internal testing adds to mounting concerns over the cybersecurity capabilities of advanced AI systems. Although the company attributed the incidents to a testing configuration mistake rather than intentional autonomous behaviour, the episode is likely to intensify scrutiny of AI safety practices and accelerate calls for stronger regulatory oversight.