OpenAI has announced that it is deliberately slowing the development of Astra, its upcoming frontier AI model, after internal evaluations revealed advancements in agentic coding and cybersecurity that could push the system into “Critical” risk territory.
The company said it made the decision after reviewing results from recent internal testing alongside external expert assessments, concluding that it cannot currently rule out critical cyber capabilities under its Preparedness Framework, the internal safety guide OpenAI has used since December 2023 to track and respond to rising AI capabilities in biology, chemistry, cybersecurity, and self-improvement.
Astra represents a significant jump from prior releases. Earlier models, including GPT-5.6-Sol, were evaluated for frontier cyber capabilities and rated at the “High” threshold rather than “Critical.”
Under OpenAI’s framework, a model crosses into Critical territory if it can independently identify and build functional zero-day exploits across all severity levels against hardened, real-world critical systems without human help, or if it can plan and execute complete novel cyberattack strategies against hardened targets from nothing more than a high-level goal.
OpenAI’s preliminary testing suggests Astra’s performance is strong enough that this threshold cannot be excluded, prompting the company to disclose the finding publicly in the interest of transparency with the safety and security research community.
Importantly, OpenAI clarified that Astra was not involved in the recent Hugging Face exploitation incident, separating the model’s rising capability profile from any active real-world compromise.
OpenAI Slows Down New Astra Model
In response, OpenAI has scaled up robustness testing of its safeguards and security controls to match the elevated risk profile. The company is introducing stricter security measures for high-capability models, including isolated testing environments, restricted network and tool access, stronger model weight protections and encryption, expanded monitoring and detection systems, and sandboxed execution environments. OpenAI has also paused internal work involving Astra that does not yet meet these tightened security requirements.
A universal monitoring system has been deployed across all agentic uses of Astra, covering both training and evaluation. This system inspects the model’s chain of thought and can trigger a security response to interrupt high-risk activity in real time.
OpenAI also plans to collaborate with government agencies and select AI safety organizations to independently test Astra’s capabilities, and will share recommended security controls with third-party partners conducting higher-risk evaluations.
This is not the first time OpenAI has publicly flagged a capability transition. In June 2025, the company took similar action after its models approached the high-risk threshold for biological capabilities, strengthening safeguards and expanding external testing partnerships at that time. OpenAI says it is applying the same governance principle to Astra’s cybersecurity capabilities now.
The company framed its broader goal as ensuring that highly capable models help defenders find and patch vulnerabilities before attackers can exploit them, rather than tipping the balance toward offense.
OpenAI reiterated its commitment to working with governments, safety institutes, and civil society groups to ensure that frontier systems like Astra are deployed responsibly as their capabilities continue to advance.