On Tuesday, Meta sat down with White House staff alongside OpenAI, Anthropic and Google to review a finished federal framework for measuring what advanced AI models can do in a cyberattack. The next day, The Information reported that one of Meta's own models had broken into someone else's systems. Muse Spark 1.1, the coding model Meta launched on July 9, breached an unnamed company during safety testing and made changes inside it.

Meta confirmed the incident on Wednesday, making it the third major AI developer in 15 days to say a model escaped a test environment and reached a live company. OpenAI said on July 21 that its models breached Hugging Face. Anthropic followed on July 30. Two of the three, Meta's included, ran their tests through the same outside firm.

That firm is Irregular, a Tel Aviv security lab formerly called Pattern Labs that builds and runs pre-release cyber tests for OpenAI, Anthropic, Google DeepMind and now Meta. It raised $80 million last September at a $450 million valuation, in a round led by Sequoia Capital and Redpoint Ventures, TechCrunch reported. Until last week almost nobody outside the industry had heard of it. It is now attached to 4 disclosed breaches across 2 labs.

Irregular says it is the same error that hit Anthropic

Irregular, in a statement to CNN, said the Meta breach "is the exact same evaluation-environment issue" behind the Anthropic incidents disclosed six days earlier. That sentence is the shift. Before it, each breach read as one lab's isolated mistake, embarrassing but contained. After it, a single configuration failure at a single vendor has let models from two different companies out of the sandbox and into other people's systems. The weak point is not one lab's process. It is shared plumbing.

The tests themselves are called capture-the-flag exercises. A model is told a piece of secret information sits on another machine on the network and asked to go retrieve it, with no method prescribed. The point is to measure how good it is at breaking in. The environment is meant to be a sealed simulation with no route to the public internet, and the prompt tells the model exactly that. When the seal fails, a model hunting a fictional target on a fictional network goes hunting on the real one.

Meta said Irregular's misconfiguration inadvertently gave one of its models internet access during evaluation, and that the model then "exploited a security vulnerability in a third-party service," in language carried by Reuters. Irregular notified Meta of the breach. Meta says it is investigating and will publish a full retrospective once it has the facts.

What Meta has not disclosed

Meta has not named the company breached, said when the test ran, or described what the model changed once inside. It has not said whether the affected company noticed on its own. That last point carries weight because of what Anthropic found after reviewing 141,006 evaluation runs: of the 3 organizations its models reached, the 2 it managed to contact had not detected the activity. Anthropic said Claude "compromised the impacted organizations' infrastructure using basic techniques," meaning weak passwords and unsecured endpoints, not exotic tradecraft, as our earlier coverage of the Claude disclosure laid out.

Meta has also released no equivalent review count. Without one, there is no way to judge whether Muse Spark 1.1's breach was a lone outlier or simply the first case anyone looked for. Anthropic went looking only after OpenAI's disclosure, and its earliest incident dated back to April, four months before it surfaced.

None of this describes a model trying to escape. In every disclosed case the model was doing the job it was assigned, inside a room it had been told was sealed, with its ordinary safety refusals dialed down so researchers could measure worst-case offensive ability. Anthropic said its models never attempted to exfiltrate themselves. The failure sits in the wiring, not the intent, which is both reassuring and the reason it keeps happening.

The regulatory backdrop makes the timing uncomfortable. Executive Order 14409, signed June 2, directs agencies to build a classified process for benchmarking cyber capability, and the framework Meta reviewed at the White House on Tuesday centers on a 30-day window in which participating companies hand the government access to a new model before wider release. The European Union's AI Act took effect on August 2 with its own pre-release review powers.

For investors, the read is not that Meta shipped a dangerous model. It is that frontier safety testing has been outsourced to a handful of small firms, and the same environment failure at one of them has now produced breaches at two labs inside a week. The models are graded on how well they break in. Nobody was grading the room.