Last week ChatGPT company OpenAI had to admit that a test version of its AI agent escaped its labs. Not only that, but it emerged that startup Hugging Face had been hacked by the rogue agent.
Now, in more news that doesn’t exactly inspire confidence, it turns out that the OpenAI agent did, in fact, hack other publicly available services thanks to finding four logins online. According to Bloomberg this included an account on the cloud platform Modal which it used as a launchpad for other attacks. Whatever the ins and outs of this, the agent went further than OpenAI disclosed. It’s difficult not to draw comparisons with Jurassic Park’s containment-gone-wrong storyline.
As per a timeline Hugging Face published on July 27 – and suggesting an incident timeline of July 9-July 13 – the agent broke into a testing sandbox “hosted on a third-party provider’s infrastructure,” which was Modal Labs. Modal’s CTO Akshat Bubna said the rogue agent had exploited some code from a customer. “Modal’s platform or isolation were not compromised in any way,” said Bubna.
In an update to its original post on July 28 OpenAI said: “No models planned for upcoming release were involved in exploiting Hugging Face.
“We have been finding a small number of cases where the models identified and used publicly exposed credentials at the account-level on other publicly-available services. This includes four accounts on four services as part of the Hugging Face incident.”
“Based on our review to date, we have not identified any other activity at the level of severity or scale of what we’ve shared related to Hugging Face, which involved a platform-level compromise.”
Hugging Face also disclosed that this was actually quite a lengthy incident of “roughly two and a half days” – it’s certainly tempting to think this was done in seconds, but the reality was a lot slower and so it raises questions how OpenAI wasn’t aware of it sooner. Hugging Face’s version of events said it was “an autonomous AI agent driven by a combination of OpenAI models ran an end-to-end intrusion against our platform: it was thousands of small, automated decisions, executed at machine speed across short-lived sandbox environments, with command-and-control staged on ordinary public web services.”
Reuters previously suggested that “OpenAI did not notice that its agent had gone haywire until well after the threat was contained and the FBI was alerted”. It’s pretty clear that OpenAI had no idea its agent had gone rogue until alerted by Hugging Face, so there are plenty of questions to answer. Hugging Face disclosed the incident on July 16.
Read the full article here
