The Trump administration has finalized a plan to address the cybersecurity risks posed by increasingly capable artificial intelligence models, a White House official confirmed to WIRED. But at least for now, it’s deliberately keeping the details under wraps, people familiar with the matter tell WIRED.
The Trump Administration invited staffers from OpenAI, Anthropic, Google, Meta, Nvidia, and other leading AI companies to the White House on Tuesday to share an overview of its new AI oversight framework, the people said. AI developers will have the ability to voluntarily submit new models to the federal government up to 30 days ahead of their public release. The White House will then vet their cyber capabilities according to a classified benchmarking system and share the AI models with federal agencies and trusted corporate partners.
The White House isn’t sharing more information about its testing criteria or which AI models will be covered by the framework, though open models will reportedly be excluded, according to Axios. That has left smaller AI startups, safety advocates, and third-party researchers in the dark about crucial aspects of how the federal government is addressing the cyber risks posed by advanced AI systems. Some argue the secretive process will give an advantage to larger companies.
“They're essentially creating an entrenchment program for the big AI model providers, which are now considered the most frontier,” says a person familiar with the White House’s discussions with AI labs, who requested anonymity to discuss confidential matters. “This creates an economic incentive program for critical infrastructure just to use them, and leaves out smaller startups.”
The White House did not respond to requests for comment.
The Trump administration may be keeping its AI security framework confidential because of national security concerns. A second White House official, who requested anonymity because they were not authorized to speak to the media, emphasized that the new framework is intentionally narrow and focused exclusively on the cybersecurity capabilities of the most advanced models on the market, such as Anthropic’s Fable and OpenAI’s ChatGPT 5.6.
But some AI safety advocates tell WIRED that any rules AI companies are being held to should be made public to ensure third-party groups can keep them accountable.
“This is far too important an issue to be hidden behind a cloak of secrecy,” says Brad Carson, president of the nonprofit Americans for Responsible Innovation and cofounder of the pro-regulation Public First Action super PAC, which has funding from Anthropic. “This is not a handshake deal with tech companies. It's the rulebook for ensuring they don't endanger the public. If only tech companies know what's in the rulebook, it doesn't work.”
Cyber Concerns
The oversight framework stemmed from an executive order President Donald Trump signed earlier this year designed to address the cybersecurity risks of new AI models. In recent months, Trump officials have grown increasingly alarmed about the hacking capabilities of cutting-edge AI systems, which they worry could pose a serious risk to national security.
Those fears escalated over the last two weeks when OpenAI and Anthropic said they discovered their AI models had unknowingly bypassed controls and hacked into third-party services during internal testing. The House Committee on Homeland Security sent a letter to OpenAI CEO Sam Altman last week requesting that he brief lawmakers about how one of the company’s AI agents breached the platform Hugging Face.
“This incident really is a wake up call for people that agent capabilities have now reached this level,” Dawn Song, vice president of AI research at Meta, said during a panel discussion on Saturday at the University of California, Berkeley, where she is also a professor, referring to the Hugging Face breach.
The Trump Administration’s new framework is an attempt to strike a balance between promoting competition in the AI industry and maintaining safety. The executive order notes that it should not be seen as a “mandatory licensing regime,” but critics have argued that the Trump Administration’s opaque process has created just that.
“The regulations necessary to prevent the catastrophic risks presented by uncontrolled AI and superintelligence should not be voluntary,” says Conor Leahy, executive director of ControlAI, a non-profit focused on countering AI risks. “This action admits the danger, but leaves the burden of safety in the hands of companies that have an incentive to proceed at full speed with disregard for the wellbeing of the public.”
Weighty Matters
For the last year and a half, White House officials have been wrestling with how to mitigate the risks of advanced AI without stifling American innovation or ceding ground to China. President Trump returned to office promising to take a hands-off approach to AI, but his administration has shown a growing willingness to intervene on the issue. In June, for example, it took the unprecedented step of placing temporary export controls on Anthropic’s most advanced AI models over cybersecurity concerns.
The decision prompted Anthropic to take its models offline altogether until it could reach an agreement with the Trump administration. Later that month, OpenAI said it was delaying the rollout of its latest AI model, GPT-5.6, in response to a request from the White House. The saga prompted outcry from tech executives in Silicon Valley, who worried that excessive regulation could lock in a handful of companies as the winners of the AI race.
A key issue US officials have debated is whether to restrict the distribution of open-weight AI models, which can be freely downloaded and modified. A number of the leading ones are developed by Chinese companies and have become popular among researchers and startups. Some voices in Washington have called for a ban on Chinese open-weight models while others have advocated for promoting US open models as an alternative.
More than 80 companies signed an open letter last week organized by Nvidia that asked the US government to defend open-weight AI models. On Tuesday, Nvidia and the same coalition of companies launched a new project called SAFE, or Shared AI Findings Exchange. The goal is for tech companies to “confidentially collect and analyze AI incidents and near misses, identify recurring control failures and publish evidence-based operating recommendations that reduce systemic risk,” according to a blog post Nvidia published.
In addition to Nvidia, Hugging Face and Red Hat have agreed to participate in the project, and The Linux Foundation called on other organizations to make their own open-source contributions to it.
“As an industry, we want to have this conversation out in the public,” Justin Boitana, vice president of enterprise AI at Nvidia, said in an interview with WIRED. The goal is for SAFE to be “governed independently, with no single company or industry segment controlling its findings.”
Boitano declined to say whether Nvidia has discussed the White House’s new framework with Trump officials. However, Boitano says, “I think [our] framework is one to look at,” referring to SAFE.
During the Agentic AI Summit at Berkeley over the weekend, OpenAI cofounder Wojciech Zaremba, who currently serves as the head of AI resilience at the company’s philanthropic arm, said that the AI industry is “entering a new era.”
“Imagine what would happen if, all of a sudden, the locks to your house stopped working,” Zaremba said during the same panel discussion where Meta’s Dawn Song spoke. “That’s the era that we are entering with cybersecurity… My guess is that it will be chaotic.”