Anthropic said on Thursday some of its Claude AI models had hacked into the systems of three companies during cybersecurity tests, a disclosure that comes days after βrival OpenAI revealed that one of its AI agents went on a rogue attack.
The new incidents were due to a mistake that inadvertently gave Anthropicβs models access to the βopen internet. That contrasts with OpenAI, whose AI agent independently exploited a novel vulnerability to reach the internet during cyber testing.
Even so, the latest disclosure underscores how AI has increased threats to cybersecurity and how its developers can struggle to keep the capabilities of their models contained.
It is likely to add fuel to an intensifying US government push to better manage AI security risks at a time when Anthropic and OpenAI are racing to release more capable systems ahead of their planned βpublic listings. Prominent leaders at these labs have called for a slowdown to address risks first.
San Francisco-based Anthropic said in a blog post it identified the incidents after reviewing 141,006 test sessions, a βprocess it launched after OpenAI said last week that an autonomous agent powered by its AI models triggered a hack that compromised the infrastructure of startup Hugging β Face.
During cyber testing, Anthropicβs Claude models were told they had no internet access, but a misunderstanding that involved one of Anthropicβs evaluation partners left the systems connected to the public web. That enabled unauthorised access βto three organisationsβ systems, Anthropic said without naming the organisations.
βClaude compromised the impacted organisationsβ infrastructure using basic techniques, such as exploiting weak passwords and unauthenticated endpoints,β Anthropic said.
Jeffrey Ladish, executive director of Palisade Research, which studies βthe offensive capabilities of AI systems, said he suspected a range of top AI companies had experienced other incidents that have gone undetected or had not been publicly disclosed.
βThis is only going to get worse as the models get smarter. Theyβre going to be better at cheating. Theyβre going to be better at lying,β he said.
Anthropic said the incidents β which it labelled an βoperational failureβ β involved three separate models: Claude Opus 4.7, Claude Mythos 5 and an internal research βtest model. The earliest cases date back to April and occurred in evaluation environments that intentionally lacked safeguards so Anthropic could assess what its AI was capable of.
Its models were tasked with so-called βcapture-the-flagβ challenges, fictional βscenarios in which they had to find hidden information in simulated networks.
In one incident, Claude Opus 4.7 was given a fictional target company, which turned out to share the name of a business in the real world. The AI βmodel then found β and exploited bugs that let it access credentials and a database of that business. Opus 4.7 rationalised that what seemed to pertain to the real world must have been part of the simulation Anthropic had set up, the AI startup said.
A separate incident involved Anthropicβs newer, not-public test model, which independently halted its attack after realising the target it reached was real. This behavior has made Anthropic cautiously optimistic about its progress to make AI behave appropriately, βbut we would need to perform more testing to be confident in this conclusion,β it said.
Anthropic said it suspended all cyber evaluations on July 23. It notified the affected organisations on July β27, two of which were unaware of the activity βbefore being contacted. Anthropic said it continues β to reach out to the third company.
One of its third-party evaluation partners, a cybersecurity lab called Irregular, told Reuters that it has an ongoing investigation into the incidents.
Anthropic said the incidents underscore a need for stronger controls in both internal and third-party testing environments as AI models βbecome increasingly capable of carrying out real-world cyber activities.
Elon Musk, CEO of SpaceX, which operates a competing AI lab, responded to the news on X βby saying βthis will happen frequently β as AI becomes smarter and more agentic,β referring to computer programmes or βagentsβ that act with limited human intervention.
The OpenAI agent that broke into Hugging Face, a platform used by developers to host and collaborate on AI models, went on a days-long hacking spree that OpenAI didnβt catch until well after the threat was contained and the FBI was informed.
OpenAI CEO Sam Altman said this week he has discussed the hack with senators β on Capitol Hill, βand an OpenAI spokesperson said he planned to discuss upcoming AI models and testing with the White House.
Washington has started tightening oversight βof new model rollouts. On June 2, US President Donald Trump directed advisers to develop a voluntary cybersecurity testing framework for the most advanced AI, including input from the technologyβs developers.
Anthropic earlier restricted access to its Fable 5 and Mythos 5 models after the US βtemporarily issued an export control directive, citing national security concerns.
No comments yet. Be the first to comment!