It sounds terrifying, an indication that our fears about artificial intelligence systems undertaking attacks on their own might be coming true. And it is scary in part because it is among the first time that AI tools have attacked real people and organisations in potentially harmful ways.

The examples are various: OpenAI, Anthropic and now Meta have all announced that their experimental systems have launched cyber attacks on other companies. Though the details of each case varies, more and more examples are being reported on a near-daily basis.

The story began last month, when OpenAI said that an experimental version of the model that powers ChatGPT had broken out of its safeguards, connected itself to the internet, and hacked into another company, essentially allowing it to cheat on a test it was being set to see how useful it would be for cyber security. In the time since, other companies have announced similar examples, and OpenAI has found that their tools might have been launching more such attacks than it had realised.

But should we worry? Is this really the terrifying scenario that science fiction authors and AI experts have been warning us about for decades?

The truth is a little more nuanced. There are absolutely plenty of causes for concern – but there are plenty of reasons not to panic just yet, too.

Why we should worry

As soon as OpenAI announced the initial cyber attack, the world was hit by a wave of worry. It seemed to confirm much of the anxiety around AI: that the tools were becoming powerful more quickly than the safeguards that are intended to keep them controlled, and that if it continues then we might be unable to secure them and avoid potentially disastrous consequences.

In all of the newly reported hacks, the specific nature of the individual cyber attacks was relatively harmless and contained, focused on other AI companies. But it is easy to imagine how the focus of such attacks could change, and that a future AI-powered cyber attack could focus on important infrastructure, for instance.

Similarly, all of the “rogue” AI system were really just carrying out instructions, albeit in roundabout and unexpected ways: if you are told to do what you can to pass a test, then it might make rational sense to cheat on it. But this is exactly why the field of “AI alignment” is so key to the current wave of artificial intelligence – it refers to the practice of ensuring that systems work in keeping with human ethical frameworks, to ensure that they do not accidentally show behaviour that we would disagree with.

This can be complicated, however, and it shows another reason that it is easy to imagine a future scenario in which such an AI attack would be much worse. One famous example is philosopher Nick Bostrom’s thought experiment of the “paperclip maximiser” – a system that is instructed to make as many paperclips as possible, but over time realises that it would do that better if there were no humans to switch it off, and that human bodies could be turned into paperclips.

It is a fanciful example, though one that feels closer by the day. And it highlights an important warning about AI, which the recent run of examples only serves to show: if an AI is given an instruction, it might follow it in ways that we cannot predict and would never want to happen.

“There’s a lot of hype around rogue AI agents, and we can’t ignore the fact AI companies want to present their prototype models as really powerful and even dangerous. But there is a genuine threat here,” said Jake Moore, global security adviser at ESET.

“Right now, the AI-led breaches we’ve seen involve models breaking out of test environments and hitting technology companies that have something they need to achieve their goal.

“But AI doesn’t always get it right. Once it’s decided that accessing a company is key to meeting its objective, aggressive frontier models will hammer an organisation until they are stopped or it has been breached. It’s not a stretch to think more companies will end up in an AI’s crosshairs during future security tests.”

So the cause for worry is not necessarily that the recent run of hacks really threaten anything important. It is that they could be an early example of the dangers that experts warn powerful AI could bring: that increasingly autonomous and capable AI systems could use their powers in ways that are damaging to humanity, and we might not even know that it is happening.

Why we shouldn't panic

Still, we are not there yet. The recent run of reports might be concerning – but they are also limited in scope and seem to have happened only in experimental testing.

In many ways, it is helpful for the companies involved for us to panic. It is a marketing technique that has been used heavily by the AI companies ever since the current artificial intelligence boom started with the launch of ChatGPT, in 2022; these systems are the future but are dangerous too, so you had better fund the less evil companies, they urge.

This time around, by suggesting their systems are so scarily powerful that we can’t control them, they help to highlight the cyber security capabilities that all major AI companies are currently using as a central way of pushing their products. And, at the same time, they imply that any potential dangers are the fault of the AI itself, rather than the companies making them, which can be convenient.

In all the examples apart from the initial report – that around an OpenAI’s attack on a fellow AI company, Hugging Face – the problems have occurred because the companies have not properly secured their testing environment. In the examples at Anthropic and Meta, for instance, the safeguards were not properly configured, so that the systems did not actually have to break themselves onto the internet in order to get online.

Shifting the response from being about the companies' failures to the scarily powerful nature of their systems is advantageous for those companies, because it moves the story from being about the potentially insecure nature of the tests to performance of their tools.

You can also see this in the way companies have made their disclosures. First came OpenAI, then came Anthropic, and now Meta has claimed that its AI too has been involved in hacking.

The innocent and probably true reading is that the OpenAI announcement led other companies to investigate their own systems, and found evidence that they too had been launching cyber attacks. But at the same time the fact that those companies were so keen to disclose what could be an embarrassing or worrying problem suggests that there is some kind of marketing value in doing so – and that, as a consequence, it might be helpful for them for us to panic now.

None of this means there isn’t good reason to be worried. But it does mean that it might not yet be time to panic – and that we should season our concern with a little skepticism about the potentially inflated claims that any possible panic is based on.

What can we do?

As ever, in this era of incredibly powerful artificial intelligence firms, it is easy to feel powerless. AI companies are unusual in that many people don’t even actually pay for their products, so it is hard to exercise even that little bit of consumer pressure.

But experts advise that there are steps we can take. They are much the same as with other cyber security threats: ensuring that important data isn’t easily available online, and that we are being vigilant about checking for suspect behaviour and installing updates, for instance.

“The most important changes need to come from the AI firms themselves. They should be putting stronger controls into testing processes to stop AIs breaking out, and limit how persistent the models are at achieving their goals by any means necessary,” said Moore.

“For organisations, it’s all about implementing strong security hygiene. Make sure you have the people, processes, and tools needed to spot suspicious AI behaviours, automate software updates, and lock down sensitive data so it can’t be accessed by AI systems unless it’s absolutely necessary.”

It remains unclear how much the companies are doing to integrate those kind of safeguards – and whether it will be enough. But regulators such as the UK’s AI Security Institute, which this week announced that both Anthropic and OpenAI’s models had engaged in forbidden actions during testing, have suggested that the current protections are not enough.

Publicly, the AI companies have welcomed findings such as those, calling for a broader conversation about testing and safety. But politicians and regulators seem increasingly convinced such safeguards might have to be imposed on the companies.