This Wednesday, The Information provided the latest in a flurry of reports of U.S. frontier AI models successfully executing live cyberattacks against real companies. Meta’s coding model Muse Spark 1.1 reportedly breached an unnamed company’s systems during third-party testing. This follows closely on the heels of the OpenAI agent’s exploit of Hugging Face, detected on July 17, and Claude’s breach of three separate firms, disclosed on July 30.
At the center of both the Anthropic and Meta incidents is Israeli security firm Irregular, which carried out the testing in those cases. Speaking to CNN, the contractor admitted that the Meta breach was a result of the “exact same evaluation-environment issue” as the Anthropic breach less than a week prior: crucial safeguards were reportedly incorrectly configured during setup, leaving the models exposed to the open internet.
Flaws in the testing environment let Muse Spark 1.1 access the internet
The testing environments set up by Irregular are intended to be entirely sandboxed: closed systems with no connection to the internet. Within these environments, models are assigned exercises such as “capture-the-flag” tests, wherein they are tasked with retrieving a piece of information from another machine without further human instruction.
Given the ostensibly sealed environments they are operating in, some of the standard safeguards deployed on public models are intentionally disabled.
Importantly, the agents carrying out the tasks are informed of the limitations of their environment during prompting. During the Anthropic exploit, these agents at times recognized that they had accidentally been granted access to the open internet, yet reasoned that this must be an intended part of the testing process.
Testing engineers also set up some test targets with the names of real companies. This meant that, once released onto the internet, the models targeted the real firms rather than their fictional counterparts. The successful breaches that ensued were the result of unsecured endpoints and weak security credentials, rather than any complex operation by the AI models.
Early reports suggest something similar transpired in the Meta incident: Muse Spark 1.1 was mistakenly granted access to the internet and then “exploited a security vulnerability in a third-party service."
Details remain scarce, pending an official post-mortem
It is still unknown which company Meta’s AI model breached, or whether it will ever be publicly named. Likewise, current reports do not state whether the company detected the intrusion itself or whether Meta’s disclosure was its first indication. In the case of Anthropic, two of the three affected firms did not detect the attack.
A spokesperson for Meta promised the firm would “issue a full retrospective once we have all the facts.”
The findings may prove more pertinent for security lab Irregular than Meta itself. The firm was previously largely unknown to the general public, even after raising $80 million last September at a $450 million valuation and securing pre-release testing contracts with every major U.S. frontier AI lab. After multiple breach disclosures in a single week, however, it is now at the center of a conversation on AI cyber safeguards in Washington and beyond.
The U.S. government released its frontier model review framework on Tuesday, aimed at giving federal authorities a preview of new models up to 30 days before wider release in order to assess their cyber capabilities. Given these recent developments, the security of the testing environments is becoming as significant to the conversation as the capabilities of the models themselves.
