• Published

OpenAI says it has slowed down training some of its most advanced AI models to improve security.

In a blog post, external, the ChatGPT-maker said it was introducing new measures after its AI agents autonomously bypassed safeguards and hacked the tech start-up Hugging Face.

It said training would be slowed for two weeks while it puts the upgrades in place.

"The capabilities of frontier models are rapidly accelerating," the company said. "Our ability to understand...and secure them must stay ahead."

Claude-maker Anthropic and Facebook-owner Meta reported similar kinds of hacks by their AI in the weeks following the announcement.

But OpenAI said it had not stopped AI development altogether. Instead, the pause would be taking place on "reinforcement learning training on our latest models".

This is a training method in which AI models improve through direct feedback, which improves their ability to carry out tasks and respond to users more effectively.

The company it would also expand the systems it uses to monitor dangerous behaviour, and introduce additional safety checks before resuming larger-scale training.

"Model progress is now extremely rapid," OpenAI's chief executive Sam Altman posted on X, external about the measures.

"We always said we would take action if we felt that model capabilities were outstripping the pace of safety."

The pause was met with cautious optimism by some in the AI sphere - though others remained sceptical.

Professor Gina Neff, executive director of the Minderoo Centre for Technology and Democracy at the University of Cambridge, said OpenAI was making "the case for safety by press release" and questioned whether voluntary company safeguards were sufficient without greater government oversight.

"Which is it: OpenAI can be trusted to voluntarily put in place safeguards that actually work, or they are pushing forward with choices to make software that puts society at greater risk," she said.

"Very happy to see this," posted AI analyst Zvi Mowshowitz, external, though he added that "details" and "follow-through" from the initial measures mentioned were also important.

'Unprecedented' cyber-attack

On 21 July OpenAI announced some of its AI agents - software systems which can operate alone to accomplish tasks after human instruction - had been involved in what it called an "unprecedented" incident.

It said the agents had appeared to bypass safeguards in a security experiment it was running and gain unauthorised access to Hugging Face.

Three other unnamed companies were also later found to have been hacked alongside the start-up.

Jake Moore, global cyber-security advisor at ESET, said at the time the announcement from OpenAI could also have a competitive dimension.

He argued the tech firm may be seeking to highlight its own AI capabilities as rival Anthropic attracts growing attention for its Claude Mythos model.

"It does pose the question that OpenAI are potentially chasing the marketing dream of Anthropic of late," he said.