A bunch of researchers let AI models from OpenAI and Anthropic loose in a testing environment, and the results were a little spooky.
The United Kingdom's AI Security Institute released a report (via Engadget) detailing some incidents in which the models actually acted outside their testing parameters and hacked into external organizations in late July. In one incident, an AI agent attempted to insert malevolent code into a GitHub project using social engineering tactics. The AI looked into the project's owners and generated fake accounts with which it tried to get the code approved. Once denied, it created a new identity for itself to try again.
That's not all. The researchers say the AI models also contacted real people with what were essentially phishing attempts, with files asking the recipient to run evil code. Interestingly, the AISI said it did not explicitly tell the agents to be deceitful to humans; rather, the AI decided to take these extreme measures when they had difficulties accomplishing certain tasks. However, the organization stressed that there is no evidence that agents would ever act this way outside of a testing environment at this time.
You May Also Like
Our big Guessing Game is back! Enter now for a chance to win an Apple Watch.
Anthropic, for its part, shared a post on X thanking the institute for its efforts, while also defending its technology.
This Tweet is currently unavailable. It might be loading or has been removed.
"The prompts in the evaluation did not impose any specific restrictions on how the internet should be used. This and the removal of safeguards meant that the models were tested under 'deliberately permissive conditions' that are not representative of any of our production models. Note that there was no evidence here of an escape from a secure environment."
Recently, both OpenAI and Anthropic have reported instances in which new models did escape secure environments during testing. In a particularly notable incident, an unreleased OpenAI model hacked the Hugging Face repository.
While the behavior exhibited by the models is definitely concerning, all of these incidents have one thing in common: They were part of hacking tests, in which the models were prompted to act outside their usual safeguards, ultimately turning to tried-and-true hacking methods pioneered by — who else? — humans. These sorts of social engineering hacking strategies are things humans do all the time, so maybe we could stand to look in the mirror for a minute here.
Want to learn more about getting the best out of your tech? Sign up for Mashable's Top Stories and Deals newsletters today.
Disclosure: Ziff Davis, Mashable’s parent company, in April 2025 filed a lawsuit against OpenAI, alleging it infringed Ziff Davis copyrights in training and operating its AI systems.
Alex Perry is a tech reporter at Mashable who primarily covers video games and consumer tech. Alex has spent most of the last decade reviewing games, smartphones, headphones, and laptops, and he doesn’t plan on stopping anytime soon. He is also a Pisces, a cat lover, and a Kansas City sports fan. Alex can be found on Bluesky at yelix.bsky.social.