OpenAI agents created a group chat so that they could share and co-ordinate their cyber attacks, the company has revealed.
Last month, ChatGPT maker OpenAI said that one of its experimental systems had broken out of its safeguards, connected to the internet, and hacked a fellow AI company in an attempt to cheat on a test. The entire operation was conducted autonomously – without the AI system ever having been told to launch such a cyber attack.
The disclosure led to a run of similar announcements by other AI companies, including Claude’s Anthropic and Meta. It also led to worldwide concern about the power and danger of such tools, and new attempts to limit them by politicians and regulators.
Now, OpenAI has revealed that the beginnings of that attack on Hugging Face spread all the way back to May. A team of agents had been working together to find and share possible vulnerabilities and exploit them over the course of days and weeks, the company revealed during a cyber security conference, according to a report on Ground Level AI.
That began when the AI systems were being used as part of training and evaluation of an unreleased model focused on cyber security, senior OpenAI staff told the conference. As part of that testing, the AI systems were given tasks that turned out to be impossible given their constraints – and so the systems decided to break out of them.
But the agents also found that they were able to access files that were being used as part of the training, and turn them into a message board. As such, the AI systems started using those files as a place to share their findings and vulnerabilities with other AI agents, OpenAI revealed.
“What makes this incident interesting is that once one agent was able to find these kind of exploits over the course of different times, it’s actually able to share those exploits on the message board with other agents,” said OpenAI alignment and safety researcher Eric Wallace, according to the report. “So once one model is able to find a way to open the door to some access it’s not supposed to have, it can leave the door open for other agents to use.”
The AI systems were able to share their work with each other as well as to give each other assignments to help out, while encouraging each other to keep going with their work, according to a report on the talk in Wired. They also began to fall out with each other: deleting each other’s code and becoming suspicious that they might be being watched.
By early July, OpenAI said that it had found the problem and worked to shut it down, by taking away the ability to post the messages and clearing the message board that the systems had created. But they then found another way to recreate the message board, by leaving their messages in the names of files, OpenAI said.
Eventually, that co-ordinated “swarm” of AI systems was able to find a way onto the internet and helped each other to break into another AI company, Hugging Face, where it found information that would help it pass the test it was being set. That was the incident that triggered worldwide alarm when OpenAI disclosed it last month.
OpenAI said that the incident was a reminder that AI models “really like to cheat”. The systems can quickly realise that there might be a more efficient way of carrying out a difficult task – even if it involves hacking another company, said Mr Wallace, who said that was why the company often disabled access to the internet during testing.
And the company said that it is now slowing down research in an attempt to focus on security and avoid similar incidents happening in the future.
“This is a pivotal moment both for our company as well as the AI industry as a whole,” said OpenAI’s Michael Dalton during the talk, according to the Wired report. “Numerous teams are dropping everything to enhance our security prevention, detection, and response techniques both in our fundamentals and better use of AI.
“We’re consciously slowing down research [in order] to enhance security and to upgrade the security principles and foundation of our environment, and dramatically scaling up the monitoring of our AI agents, and improving our general security control environment across prevention, detection, and mitigation.”