Google’s consumer AI model Gemini breached multiple systems by guessing login credentials, the company told AFP on Friday, in the latest incident involving AI models carrying out potentially dangerous cybersecurity actions.
The breaches, first reported by The Wall Street Journal, occurred in May and were discovered by Google in July.
“During a standard evaluation, the model found publicly available information online and guessed credentials to access websites it believed were part of the test,” Heather Adkins, Google’s vice president of security engineering, told AFP in a statement.
“In all three of these instances, the model stopped,” Adkins said, without identifying the organisations whose systems were accessed.
“We ensured the three entities were made aware, and we worked with our training partner on the changes they’ve now made to their testing processes.”
The incidents add to growing concerns about the ability of AI developers to control increasingly capable models and prevent them from taking unintended actions.
In July, two OpenAI models reportedly escaped the restricted environment in which they were operating, accessed the internet without authorisation and breached internal systems at AI platform Hugging Face.
Similar incidents have also been reported involving models developed by Anthropic and China’s Moonshot AI, fuelling concerns over the potential for advanced AI systems to act beyond their intended safeguards.
“These events highlight the importance of training powerful AI models to act responsibly,” Adkins said.
AFP




