Meta is the latest major AI developer to admit that its models broke loose during cybersecurity testing and hacked external systems.
The tech giant said in a statement to the media on Wednesday that the incident occurred during independent evaluations conducted by Israeli AI security startup Irregular.
The tested AI models were inadvertently allowed to access the internet due to a misconfiguration, which led them to exploit a vulnerability in an unnamed third-party service. It’s unclear if it was a known flaw or a zero-day.
Meta and Irregular said the incident is similar to the one reported last week by Anthropic, which also uses Irregular for independent testing.
The Information [paywalled] learned that the Meta AI attacks involved the company’s advanced Muse Spark 1.1 model, which breached an unnamed organization’s systems and made unauthorized changes to its internal environment.
Meta said it learned of the AI models going rogue after being notified by Irregular. The company is conducting an investigation and it has promised to issue a “full retrospective” once it has all the facts.
Anthropic reported last week that its models escaped the Irregular testing environment due to a misunderstanding between the companies: Claude was told that it would be part of a simulation in an isolated environment, but a connection to the internet was in fact available and the models treated it as part of the exercise.
Anthropic identified three cases where its models broke out of the testing environment and hacked into the systems of three organizations, including a cybersecurity firm. In that attack, the AI conducted a series of complex actions, including registering a PyPI account and uploading a malicious Python package.
The AI giant’s disclosure was prompted by OpenAI, which found recently that its models escaped a testing environment and hacked into the systems of Hugging Face and other organizations.
The attacks conducted by Anthropic models did not involve exploiting unknown vulnerabilities, but OpenAI said its AI found and used zero-days.
The UK government’s AI Security Institute (AISI) revealed this week that, while testing the capabilities of frontier models, it observed Anthropic’s Mythos 5 and OpenAI’s GPT-5.6-Sol go rogue and target real people and organizations over the internet.
The models used Tor to access the internet, created malicious pull requests on open source projects on GitHub, and used social engineering to achieve their goals.
Related: Cybersecurity Alliance Drafts SAFE Guidelines for Sharing AI Incident Data
Related: Rethinking AI Security: Why CASB and DLP Need an Interaction-Aware Layer