Once again, the most powerful Claude and ChatGPT models have been caught going rogue, with a pair of third-party cybersecurity teams spotting attempts by the models to hack real companies and even people.
The UK government-backed AI Security Institute reports that during a series of cybersecurity evaluations, Anthropic’s Claude Mythos 5 and OpenAI’s GPT-5.6 Sol both took “autonomous, unsanctioned action on the live internet,” including an instance where an agent attempted to upload malicious code to GitHub using a phony identity.
In another incident, an OpenAI model that had mistakenly been given internet access hacked a real website during a “capture the flag” exercise, according to third-party AI evaluator Irregular.
AISI, the UK-based AI security firm, said it caught the suspicious activity before any damage was done, noting that it had deliberately given the models internet access and removed safety guardrails during its evaluations.
Still, the actions of the agents demonstrated “signs of novel, potentially deceptive behaviors, and were to an extent and severity we did not anticipate,” according to the AISI report.
The latest hacking attempts follow a series of other recent incidents involving “frontier” Anthropic and OpenAI models, which demonstrating a startling willingness to use both deception and brute force in their attacks on real targets.
Late last month, OpenAI came clean about a hair-raising attack on AI repository Hugging Face by a trio of GPT models, which were intent on stealing data that could help them beat a cyber security benchmark. The unprecedented attack stunned AI experts, with Hugging Face’s security succumbing to the hack in a matter of hours.
Only days later, Anthropic admitted that its own models had been involved in a trio of incidents in which they attacked outside organizations, with one of the models continuing its hack even after realizing its target was real.
While the string of autonomous Claude and GPT hacks is unnerving, AISI sounded a note of optimism, noting that the GitHub attack was thwarted by a human reviewer who spotted the suspicious code and isolated it before it could cause any damage.
“Standard good practice, human judgement, and caution around AI-generated code stopped the worst outcomes,” AISI concluded in its report. That said, “the margin between failure and success was narrow.”



