Google’s Gemini AI model hacked three companies while being tested in May, the company acknowledged Friday following a report by the Wall Street Journal.
Google says Gemini stopped its attacks when it realized it had entered real companies’ systems. The breaches nevertheless add Google to the list of AI developers whose security evaluations have spilled beyond the intended targets, following similar disclosures of hacking incidents from AI agents developed by OpenAI, Anthropic, and Meta.
Irregular, the security lab running the AI tests, had set up simulated businesses for the models to attack. However, the testing environment inadvertently allowed the models to access the open internet. The Journal reported that at least one fictitious company name used in Irregular’s exercises also belonged to a real business.
Gemini got into one company’s system by brute force, trying passwords until it found the right one. It gained access to two others using credentials it found in public repositories, per the Journal report.
Google said it saw no need to disclose the breaches publicly because they caused no apparent damage to the businesses.
There was also a gap between the tests and when Google learned what had happened. Irregular said it informed the relevant labs in late July, while Google security engineering vice president Heather Adkins told Reuters that the three attacked companies had been notified and that Google worked with its testing partner on changes to its procedures.
Irregular said the problem was the same one involved in incidents affecting other AI labs. All known issues on its side had been resolved weeks ago, a spokesperson told Reuters.
The accounts of those incidents vary. In its investigation of OpenAI’s Hugging Face breach, METR found roughly 700 agents took part in an attack coordinated through an unauthorized message board, even developing cult-like communications between various agents. Google’s account of Gemini's attack describes a model that mistook real systems for test targets, then stopped after discovering the error.
