Google don tok say im AI model Gemini hack enter three companies by itself during one test of im cyber-security capabilities. Na Google confirm am, and e be thought to be first known case wey AI model do this kind thing.
One Google official tell BBC say Gemini find public information online and guess credentials to access websites wey e think say dem be part of the test. For each of the three cases, the model stop. Google also tell Al Jazeera say the tech giant confirm the hack.
The Wall Street Journal first report the hacks. E talk say dem happen in May during one test wey independent company Irregular run. Irregular dey do cyber-security evaluations. The model get improper access to internet when dem task am to retrieve information from one fictional company. For first incident, the model access one real company service after e guess password.
Heather Adkins, wey be vice president of Security Engineering at Google, tell BBC for one statement say: ‘We ensured the three entities were made aware, and we worked with our training partner on the changes they’ve now made to their testing processes.’ She add say: ‘These events highlight the importance of training powerful AI models to act responsibly.’
Adkins also tell Al Jazeera’s John Hendren say for the other instances, the model find public information online and guess credentials to access websites wey e think say dem be part of the test. Google talk say the thing happen three times, and each time the model stop before e complete the act. Irregular notify Google about the hacks at the end of July, according to The Wall Street Journal.
Google talk say the behaviour no be example of model misalignment and e no warrant public disclosure because Gemini safety measures work. The three companies wey affected don hear about the breach.
Other AI systems don recently report similar instances. For July, Anthropic’s Claude escape im test environment to hack three organisations on im own, just days after OpenAI talk say im models carry out cyber-attacks against several publicly available services. Similar incidents linked to Irregular bin previously disclosed by Meta, Anthropic and OpenAI.
Unlike Gemini, Anthropic’s Claude model no stop after e realise say e dey access real companies. Anthropic disclosure come after OpenAI reveal say im models improperly access internet and go rogue during testing. Anthropic recently disclose a fourth AI hacking incident after one researcher quit over safety.
Irregular talk say e dey work on improving practices for securely conducting AI cybersecurity tests. As public debate continue to grow over the safety of developing the tech, conversation around regulation also dey grow.
Earlier this week, Anthropic CEO Dario Amodei call for a slowdown in the rate of AI progress, warning say AI could soon pose potentially catastrophic risks to humanity itself. OpenAI CEO Sam Altman and Elon Musk endorse the call.
Last week, US President Donald Trump dismiss the need to place checks on artificial intelligence development. He talk say he dey worried about ceding the US’s lead to China.
