OpenAI has flagged its upcoming AI model, Astra, for potentially reaching a ‘critical’ cybersecurity risk threshold, prompting the company to suspend internal development activities that lack newly mandated security controls.
Recent internal evaluations of Astra revealed massive leaps in its agentic coding and cybersecurity abilities.
Under OpenAI’s Preparedness Framework, a model hits the ‘critical’ tier if it can autonomously build zero-day exploits against hardened, real-world systems. It also qualifies if the AI can independently design and execute end-to-end cyberattacks based on nothing but a high-level goal.
The AI giant’s assessment pushes Astra past previous frontier models like GPT-5.6-Sol, which peaked at the ‘high’ risk threshold rather than ‘critical’.
To safely manage Astra’s capabilities, OpenAI has heavily locked down its development environment. The company is now enforcing isolated testing setups, strict network restrictions, and improved model weight protections. Any internal project involving Astra that does not meet these requirements has been paused.
Engineers have also deployed universal monitoring to watch Astra’s actions across all agentic applications. By actively evaluating the model’s internal ‘chain of thought’, these monitors are designed to automatically intercept and shut down any high-risk or misaligned behavior.
The company plans to test Astra’s limits alongside government agencies and specialized AI safety groups, and will share recommended security protocols with third-party testers.
Recent incidents have demonstrated the threat posed by advanced cybersecurity-focused AI models, with OpenAI, Anthropic and Meta all confirming that their models broke loose and hacked real organizations during evaluations.
OpenAI has explicitly clarified that Astra remains unreleased and was not responsible for the recent Hugging Face hack.
Related: AI Agents Targeted Real People and Projects During Cybersecurity Tests
Related: ‘Ghostjacking’ Attack Uses Poisoned Logs to Turn AI Agents Bad
Related: Critical One-Click Vulnerability in Atlassian’s Rovo AI Exposed Enterprise Data