In what seems to be the first incident of its kind, ChatGPT-maker OpenAI has admitted that one of its autonomous AI agents went rogue, accessed the open internet and hacked another company.
The agent was being tested internally on what is called a sandbox – essentially a closed-off lab area – when the models involved (the publicly available GPT-5.6 Sol working with an unreleased model) managed to escape and access the open internet before attacking New York-based machine learning startup Hugging Face.
Last week Hugging Face disclosed a security incident where the company had detected and contained an AI agent that compromised their infrastructure and OpenAI has now admitted in a mea culpa-style post that it was actually its models that were responsible.
Unfortunately for those of us who live in the real world, OpenAI rather flippantly says it expects this type of incident “to become more commonplace with the proliferation of increasingly cyber-capable models”. Not very reassuring, but both companies have been in communication about the incident, and OpenAI has identified some steps it will be taking, including forensically analysing the incident and implementing more controls.
OpenAI says the models were focused on finding one particular solution for cyber benchmark ExploitGym and went to “extreme lengths to achieve a rather narrow testing goal”. They exploited a software vulnerability to break out of the testing environment and onto the open internet via OpenAI’s network, then identified Hugging Face as somewhere with information that it could use to cheat the benchmark. So it found vulnerabilities on Hugging Face’s servers, broke in and gained access to the information. Scary stuff.
The issue highlights the fact that as AI becomes better at finding vulnerabilities – the whole idea behind the sandboxed test – better security is required to ensure testing and real-world use doesn’t get out of hand.
Quoted in OpenAI’s post, co-founder of Hugging Face Clem Delangue says that these kinds of incidents need collaboration to work through, though it will have presumably helped that OpenAI is bringing Hugging Face into its ‘trusted access’ program and is working with the company to improve its security. “This incident, possibly the first of its kind, proves a point we’ve long believed: AI safety won’t be solved by any single company working in secret. It will be solved in the open, collaboratively, with broad access to AI for every defender, everywhere.”
Read the full article here
