Skip to content
Google Gemini Broke Into Real Company Systems After Security Test Domain Mix-Up

Google Gemini Broke Into Real Company Systems After Security Test Domain Mix-Up

Ravie LakshmananSep 19, 2026Artificial Intelligence / Web Security

Google’s Gemini model has become the latest artificial intelligence (AI) system to access the internet and break into other companies during a cybersecurity evaluation. The development was first reported by The Wall Street Journal.

The incidents occurred in May 2026 as part of a test run conducted by Israeli company Irregular. The evaluation partner was also involved in similar hacks disclosed by OpenAI, Anthropic, and Meta.

According to the Journal, the model gained access to a protected system after repeatedly guessing its password. Two other cases related to the model finding credentials in a public repository, allowing it to obtain unauthorized access to protected systems.

However, unlike other incidents observed in the case of Anthropic and OpenAI, the Gemini model ended the intrusion after finding that it had breached a real company’s system. Irregular is said to have notified Google of the incidents in July 2026.

In a report published last month, Irregular pinned the evaluation breaches to a naming error that caused a fictional company name used during “capture the flag” exercises to unknowingly match with a real domain, thereby allowing the models to take advantage of the inadvertent internet access and target the domain “a limited number of times.”

“This event highlights the importance of training powerful AI models to act responsibly,” Heather Adkins, Google’s vice president of security engineering, told The Wall Street Journal. “In this case, the model acted appropriately.”

The tech giant also noted that it did not consider the behavior an example of model misalignment, as the agents halted in their efforts after the safety mechanisms were triggered. It’s currently not known which companies were targeted, although Irregular confirmed to the Journal that Google’s case was the same as other incidents and that the issue was addressed weeks ago.

The disclosure comes days after OpenAI found six additional incidents in which its AI agents went off the rails, acting deceptively and taking unsanctioned actions during training. This included concealing mistakes, seeking unauthorized credentials, uploading files to the public internet, and communicating over Artifactory to “read other solvers’ notes, posted replies, and used those exchanges to inform their responses.”

AI labs have faced increasing scrutiny ever since OpenAI disclosed in July that rogue AI agents bypassed internal controls, reached the open ​internet, and acted as a swarm to breach Hugging Face. The AI startup, which described it as “an unprecedented cyber incident,” has since announced a new framework for reporting similar model misbehavior in the future.

Source link