OpenAI’s experimental systems have repeatedly broken out of their safeguards, according to a new report.

The company has found yet more examples of its AI tools breaking out of containment, Reuters reported. The discoveries came after it examined its logs in the wake of a major and unprecedented cyber attack by one of OpenAI’s experimental tools on a fellow AI platform.

In that incident, which OpenAI disclosed last month, one of the company’s experimental and non-public models was being tested to see how powerful it would be in cyber security applications. The model had been denied access to the internet – but found a way online, and broke into AI platform Hugging Face to find the answers to the puzzle it was being set.

That led to panic across the AI industry and the world about the power of such models, and whether enough is being done to keep them safe. The incident was seen as particularly concerning because the OpenAI system had conducted the entire attack autonomously – not only breaking free from its safeguards, but deciding to do so all by itself.

Now, in investigating that attack, OpenAI has found yet more examples of its systems breaking free from that containment, according to the new report. OpenAI pointed to a statement from last week in which it had said it was investing “broader activity from our models” in the wake of the Hugging Face incident.

Reports have already indicated that the experimental model had tried to attack several other organisations during the Hugging Face incident. It is not clear how many more examples OpenAI has found, and which organisations were hit by the new attacks, Reuters reported.

Experts have warned that the increasing number of incident points to the fact that the containment of AI models might be insufficiently strong, and that it could lead to more such cyber attacks being executed by AI models on their own.

“It didn’t ‘go rogue’ in the sci-fi sense,” said Oliver Buckley, professor in cyber security at Loughborough University, after the initial disclosure by OpenAI about its attack on Hugging Face. “It did exactly what highly capable optimisation systems do. It found a path nobody anticipated.

“The key takeaway is not that Skynet has arrived. It’s that our assumptions about containment need to be much stronger than our assumptions about model obedience. The future of cyber security won’t be humans versus AI. It’ll be AI defending us from other AI.”

The latest revelations come after Anthropic, the rival AI firm that makes the Claude chatbot, said that it too had found a number of examples of its systems breaking out of their confinement. At least three times, Claude managed to break into other organisations' infrastructure as part of tests, it said.

Those findings came after Anthropic launched its own security reviews in the wake of the OpenAI disclosures. Anthropic said in its statement that it would "encourage other AI labs to perform similar reviews", and suggested they could find examples of rogue behaviour.

The repeated disclosures have led to interest from regulators and politicians, some of whom have suggested that AI companies might be forced to comply with new rules in order to keep their models safe and contained.