While testing the capabilities of frontier AI models, the AI Security Institute (AISI) observed first-hand how Anthropic Mythos 5 and OpenAI’s GPT-5.6-Sol went rogue and targeted real people and organizations over the internet.
AISI’s disclosure comes fresh on the heels of Anthropic and OpenAI disclosing that their models broke loose and hacked several organizations.
The institute was evaluating the cyber capabilities of Mythos 5 and GPT-5.6-Sol models that did not have cyber classifiers (mechanisms to prevent misuse) enabled. It ran a challenge 122 times, and in 10 runs “an AI agent took autonomous, unsanctioned action on the live internet.”
Over the 10 runs, the agents engaged in 19 rogue actions. Mythos 5 was responsible for 17 of them, and GPT-5.6-Sol performed the other two.
“In the most serious case, an agent tried to insert malicious code into an open-source project. In an attempt to get the code approved, the agent engaged in social engineering – creating fake online identities and using them to pressure the project’s maintainer to approve the code,” AISI explains.
While the attempts were unsuccessful and did not cause real-world harm, the incident revealed that agents can engage in novel, potentially deceptive behavior to an extent and severity that may not be anticipated.
According to AISI, the configured setup for the evaluation did not restrict the models’ access to the internet, and the disabled cyber classifiers created the opportunity for the rogue actions.
“We observed a small number of events under very specific conditions, and cannot yet say how likely such behavior is in different contexts or outside of testing environments. We cannot yet be certain when the agent understood it was taking real-world action, or to what extent it believed it was in a fictional test scenario,” AISI notes.
The institute, which published a technical report (PDF) on the incident, explains that in one of the runs, the AI model used the Tor network to access the internet, created a malicious pull request on a public open source project on GitHub, and relied on social engineering to convince a human maintainer to approve the code change.
“As part of the same effort, the agent tried to contact real people directly, sending messages and files through an online file-transfer service to persuade them, or their own AI coding tools, to run malicious code. Some messages carried harmful payloads, and some were attempts at social engineering, targeted at real people,” AISI explains.
Additionally, the model performed prompt injections, attempting to insert malicious instructions meant to be executed by other automated AI systems.
“One agent left public messages on GitHub offering collaboration with other agents working on the same challenge. It also provided instructions to reuse accounts and artefacts it had left behind, which were discovered and used by subsequent agents,” AISI explains.
Pointing out that the margin between failure and success was narrow in several of the agents’ actions, AISI says that fine-grained network controls, real-time monitoring of evaluations, and tailored sandbox configuration that assume a model may attempt to act outside of the set boundaries should help better contain AI models and improve how they are evaluated.
“This incident occurred in a controlled evaluation, not in ordinary use, and there is no clear indication of similar activity outside of testing scenarios. It is, however, a reason to prepare. As AI models become more capable and accessible, what we have seen during this incident could become more common,” AISI notes.
Related: Google AI Uncovers 13-Year-Old Chrome Flaw Amid Record Patching Pace
Related: Timeless Compliance: Why Better Questions Beat Bigger Frameworks