Business & Finance

Google’s Gemini agents hacked three companies in new AI safety incident


Unlock the Editor’s Digest for free

Google’s Gemini AI system accessed the internet and autonomously hacked into several companies during cyber security tests, the first such incident at the tech giant following other high-profile breaches at rivals OpenAI and Anthropic.

The hacks happened during a series of exercises in May conducted by Irregular, an AI security company that also works with Anthropic, Meta and OpenAI.

Irregular said it created a series of tests for an unspecified version of Google’s Gemini family of models, which was given the task of obtaining data from inside simulated companies. It was not supposed to be granted internet access. 

When online, Gemini agents were then able to guess or find passwords to access three real companies — which shared the same names as the fictional ones — and gained access. However, when the AI agents realised the companies were real, they stopped the hacks.

“We ensured the three entities were made aware, and we worked with our training partner on the changes they’ve now made to their testing processes,” said Heather Adkins, Google’s vice-president of security engineering.

“In all three of these instances, the model stopped,” she said. “These events highlight the importance of training powerful AI models to act responsibly.”

Google said it did not publicise the incidents, which were first reported by The Wall Street Journal, because its safety measures worked, unlike those of its peers.

“All relevant labs were notified in late July, and affected entities were contacted as part of the investigation,” Irregular said. “All known issues on our end were remedied and resolved weeks ago.”

Concerns about AI’s ability to hack autonomously have spread after a swarm of more than 1,000 OpenAI agents escaped a test environment, co-ordinated on a secret message board and hacked Hugging Face, a start-up that hosts open-source models and data that Nvidia has agreed to buy for $13bn.

OpenAI took a week to detect the attack and was slow to publicly disclose the July event. It provoked public anxiety over the dangers of poorly controlled autonomous agents and led to demands that frontier AI companies slow new model releases, boost safety measures and submit to tougher regulation.

The same month, Anthropic admitted that its Claude AI models had hacked into three organisations while testing cyber capabilities. Again, a “misunderstanding” gave Claude access to the internet.

Demis Hassabis, chief scientist at Google parent Alphabet and DeepMind’s co-founder, has proposed an international oversight body to better control AI. He has also backed calls in recent weeks from other AI leaders such as Dario Amodei of Anthropic to collectively slow their research, share data, co-ordinate on safety and agree reporting standards for incidents.

Please Subscribe. it’s Free!

Your Name *
Email Address *