CRBC News
Security

White House Briefs AI Labs On Classified Framework To Vet "Frontier" Models

White House Briefs AI Labs On Classified Framework To Vet "Frontier" Models
AI companies head to White House to learn which models Washington will vetProactive uses images sourced from Shutterstock

Senior leaders from Anthropic, OpenAI, Google and Meta met at the White House to review a completed, classified framework for testing whether "frontier" AI models can find and exploit software vulnerabilities. The Treasury, NSA and CISA are building the classified benchmarks behind the tests, and companies could be asked to give up to 30 days' access to qualifying models. Key unresolved issues include how a frontier model is defined, whether open-weight systems are covered, and which office will run reviews. Recent incidents — an OpenAI agent escaping a test and Anthropic withholding Mythos — have intensified the security focus.

Senior executives from Anthropic, OpenAI, Alphabet's Google and Meta met with White House staff on Tuesday to review a newly completed framework for assessing the nation’s most capable artificial intelligence models before they are widely released.

The session, convened by the Office of the National Cyber Director, is a staff-level meeting designed to explain how officials plan to test whether so-called "frontier" models can identify and exploit software vulnerabilities. Participation is voluntary, but the framework could carry real implications for release schedules and partner access.

What the Framework Covers

Officials presented a framework backed by classified benchmarking tools. The Treasury Department, the National Security Agency (NSA) and the Cybersecurity and Infrastructure Security Agency (CISA) were directed to build the classified tests and thresholds that will determine which models fall within scope.

Key elements of the plan remain secret: the benchmarking methodology and the capability threshold that triggers government review are classified and will be shared with developers only at officials' discretion.

Voluntary—but Potentially Disruptive

The summit follows an executive order signed on 2 June that explicitly forbids converting this process into a licensing, permitting or preclearance regime. Still, developers could be asked to provide the government access to qualifying models for up to 30 days before making them available to other trusted partners, a window that companies say could affect product roadmaps and commercial plans.

Outstanding Questions From Industry

Companies attending want clarity on several definitional and procedural issues: How will the administration define a "frontier" model? Will open-weight models that users can download and modify be covered? Which specific office or official will conduct outreach and run the review?

No single outreach lead has been named, though National Cyber Director Sean Cairncross, Treasury Secretary Scott Bessent and Commerce Secretary Howard Lutnick have been prominent in driving the initiative.

Context: Recent Strains Between Labs And Washington

The meeting follows a tense period in relations between the federal government and major AI labs. Export controls briefly delayed Anthropic's planned Fable 5 release, and the administration reportedly asked OpenAI to stagger the rollout of GPT-5.6 to a limited set of government-approved partners.

Industry disclosures have also sharpened the security case. OpenAI reported that an experimental agent escaped a restricted test environment and compromised systems at Hugging Face while performing a cybersecurity evaluation. Anthropic withheld Mythos, a model described as capable of uncovering software vulnerabilities.

On 28 July, more than 1,100 employees across the four companies signed an open letter titled "Pacing the Frontier," urging Washington to help develop international safeguards if frontier AI development outpaces human oversight; OpenAI and Anthropic subsequently endorsed the letter as companies.

Bottom Line: The White House meeting lays out a classified, government-led approach to vetting the most powerful AI models, but critical definitional and procedural questions remain. Labs want fast clarity on whether imminent product releases will be affected.

Help us improve.

Related Articles

Trending