OpenAI: Upcoming Astra Model May Reach Critical Cyber Capability Threshold

OpenAI reports that internal evaluations of Astra, an upcoming model, show sufficient performance in agentic coding and cybersecurity that the company cannot rule out that it meets the Critical threshold under its Preparedness Framework.

The Critical cybersecurity threshold is defined as a model that can identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention, or can devise and execute end-to-end novel strategies for cyberattacks against hardened targets given only a high-level goal.

Previous models, such as GPT-5.6-Sol, were assessed at the High threshold.

In response, OpenAI has implemented stricter security controls including isolated testing environments, restricted network and tool access, enhanced model weight protections, sandboxed execution, and universal monitoring for risky actions across all agentic applications of Astra.

Internal activities involving Astra that do not yet meet these strengthened controls have been paused.

OpenAI will work with relevant government agencies and select AI safety organizations to test the model’s capabilities and will provide recommended security controls to third-party testing partners.

The company states that Astra was not involved in exploiting Hugging Face and that it is sharing this information for transparency with the public and safety communities.

The Preparedness Framework, first published in December 2023, previously guided similar actions in June 2025 when models approached the high capability threshold for biology.

Responding to the next frontier of critical cyber capabilities

View Original