The U.S. federal government is finalizing its framework for assessing frontier AI models’ cyber capabilities this week, with industry leaders and White House officials set to convene on Tuesday to discuss its contents. Central to the framework is an optional review window, with participating developers submitting their new models to government agencies up to 30 days before wider release.

This comes less than three weeks after OpenAI announced that one of its models “broke containment” during testing and compromised open-source AI platform Hugging Face, becoming the first publicly documented breach of its kind. As OpenAI noted in its post-mortem report, "the containment architecture itself became part of the attack surface."

This incident brought the question of cybersecurity safeguards for AI models into sharp focus, particularly the guardrails set up during internal testing. Anthropic later disclosed that its own models managed to breach systems at three firms after sandbox guardrails were improperly configured when setting up the tests.

OpenAI has already been offering federal officials previews of its unreleased Astra models in recent weeks, ahead of the framework’s finalization. It is expected to attend the White House meeting alongside representatives from Anthropic, Google and Meta.

How lawmakers have responded to AI cyber threats

In the U.S., lawmakers are pushing to address potential AI threats at both the state and federal levels. California, New York and Illinois have enacted mandatory incident disclosures for frontier AI developers. Fifteen Republican state attorneys general are also reportedly considering taking OpenAI to court over consumer protection claims.

Executive Order 14409 - the directive outlining this new federal review framework - was signed on June 2, in response to Anthropic’s claims about the cyber capabilities of its Mythos model. Representatives from the frontier AI labs have reportedly had access to draft review copies of the framework, with the chance to submit revision requests.

Although its exact contents are not public, the order directs federal agencies to set up benchmarking tests for new models and establish clear guidelines for which qualify as “covered frontier models.” It will also establish an industry-government clearinghouse for finding and patching software flaws with AI assistance.

Such government access to frontier models is not unprecedented; under the European Union’s AI Act, the European Commission can require developers to provide access to general-purpose AI models for regulatory evaluations. The U.S. equivalent makes participation explicitly voluntary and states that it is not a licensing requirement for new model launches.