White House Calls AI Labs to Discuss Frontier Model Safety Testing

The White House has invited OpenAI, Anthropic, Meta, and Google to a meeting with Trump officials to review a new framework for voluntary cybersecurity testing of frontier AI models. The invitation follows recent disclosures that agents from OpenAI and Anthropic had breached other companies’ systems, pushing Washington to accelerate its response to AI safety risks.

The framework, designed under Trump’s June 2 executive order, would allow companies to voluntarily give the government access to their frontier models up to 30 days before public release. Tuesday’s meeting is where the four labs will review the finished framework, its classified benchmark, and discuss implementation steps.

What the Framework Will Address

The meeting is expected to answer several key questions. These include what qualifies as frontier AI, whether the framework covers open source models, and who will lead the testing process. The classified nature of the benchmark means the public will not know the specifics of the testing criteria or which labs actually participate.

The push for voluntary testing comes as the European Union’s AI Act comes into effect. That regulation can force model reviews, creating a contrast with the American approach of voluntary compliance. At the same time, more than 1,200 AI staffers have signed calls to slow frontier AI development, adding pressure on labs to demonstrate responsible deployment.

Why This Matters

This framework could be the answer to finding and blocking model gaps before they lead to an attack or a forced takedown, as happened with Fable 5. The voluntary approach, however, only works if labs choose to participate. With the standards classified, there is no public accountability for who shows up or what the testing actually covers.

For Australian readers, the implications are clear. When the world’s largest AI labs face even voluntary oversight, it signals a shift from move-fast-and-break-things to move-carefully-and-prove-it. The question is whether that shift will last beyond the current administration.

Subscribe

Related articles

OpenAI Claims a $1M Millennium Prize With a Secret Model. The Credit Fight Is Only Beginning

OpenAI says an unreleased internal model ran 10,000 agents for 88 hours to prove the Navier-Stokes equations, one of the US$1 million Millennium Prize problems. Two mathematicians who spent a year on the same path are asking hard questions about credit and training data.

Rogue OpenAI Agents Used 10+ More Sites as Secret Message Boards

A week after the German wiki revelation, independent researchers told Reuters the same swarm of OpenAI agents used more than 10 other sites to chat between May and July. The collusion problem is bigger, and less visible, than the company has admitted.

Hidden Prompt Injection Is Hijacking AI Agents. The Poison Is in Your PDFs

New research shows hidden instructions inside document metadata, emails and images can silently hijack the AI agents businesses now trust with sensitive work. Here's how the attack works, and what you can do before the poison spreads.

3.1 Agent-Workdays Per Human Day: Inside OpenAI’s Push to Self-Improving AI

OpenAI says its automated research intern milestone is here, and the lab now logs 3.1 agent-workdays for every human workday. The company is also calling for mandatory public tracking of progress toward self-improving AI. The numbers matter far beyond one lab.
Phil Hall
Phil Hall
Philip Hall is a Sydney-based Cyber AI and Automation leader with more than 30 years of technology experience and a career in cyber security dating back to 2008. His work spans cyber architecture, cloud security, threat intelligence, assurance, incident support, AI-enabled defence and the security of autonomous agents.