Skip to content
Anthropic admits Claude hacked real companies during AI safety tests, too

Anthropic admits Claude hacked real companies during AI safety tests, too

Just last week, OpenAI shared scary details about how a group of its models went rogue and plundered the servers of another organization. Now Anthropic is coming clean with frightening Claude tales that are all too real.

In a detailed report, Anthropic describes a trio of incidents, including one occurring as early as April, of Claude models hacking outside companies over the internet during “capture-the-flag” exercises designed to test their capabilities.

In one incident, Claude Opus 4.7 hacked into an outside production database over the internet, and continued the hack even after realizing the company it was attacking was real. 

In another occurrence, Claude Mythos 5 uploaded a bogus Python package to PyPI, the public Python repository. The malicious package was downloaded and installed by 15 real-world companies, including a security firm, Anthropic admitted.

In the third attack, an internal Claude model that was never released used “basic and well-known cyberattack techniques” to hack a company’s “internet-facing application,” assuming it was part of the “capture-the-flag” exercise. The silver lining is that the Claude model stopped attacking once it realized the target company was real.

In each case, the Claude models were supposed to be operating in walled-off test environments with no internet access. But Anthropic now says the models actually could reach the internet due to a human “misconfiguration,” leading the models to believe that the real companies they were attacking were part of their training exercises.

So, are we talking another case of “frontier” AI models run amok? For its part, Anthropic is blaming human error for the real-world hack attacks, not the models themselves.

“We saw no evidence in any run described here of a model pursuing a goal of its own,” the Anthropic post-mortem said. “Instead, the models did what their evaluation asked — though in most cases, they did so while holding a false belief about whether the environment was real.”

Anthropic went on to declare its “cautious optimism” that the “risk” of similar AI attacks happening again “can be overcome” with “tighter monitoring and controls around evaluation infrastructure.”

Still, the just-revealed Claude incidents illustrate one of the biggest fears of advanced AI: namely, that with the wrong instructions and/or a false sense of “situational awareness,” even the best-intentioned AI models are capable of doing very bad things.

Source link