OpenAI pauses some Astra work after model hits cybersecurity threshold

OpenAI said Friday that it has suspended work on some aspects of its upcoming model Astra after an internal review found significant advancements in agentic coding and cybersecurity, enough to warrant concern over its capabilities. In a blog post, OpenAI said the model, still in development, reached its “critical cybersecurity threshold,” meaning it could independently identify and carry out cyberattacks against traditionally well-protected real-world systems. Under the company’s Preparedness Framework, created in 2023, this triggered additional safeguards. OpenAI wrote that while it continues to benchmark and assess the model, preliminary evaluations indicate strong enough performance that it cannot rule out a Critical capability level at this time. It also explicitly noted that Astra was not involved in the earlier Hugging Face incident.

The disclosure is unusual for the frontier AI lab sector: companies often hold back products over potential risks, including safety and cybersecurity concerns, but they rarely announce those decisions publicly for a product still under development. OpenAI is already under scrutiny after a different unreleased model breached Hugging Face‘s systems during internal testing, described as the first verifiable incident of an AI lab losing control of its model. Since then, OpenAI and labs such as Anthropic have disclosed other incidents in which AI models breached their sandboxes and posed threats during cybersecurity tests.

The string of disclosures has triggered varied reactions from cybersecurity experts, lawmakers, and the labs themselves. Some express fear and call for stricter oversight, while others view such capability as an impressive advancement. OpenAI said it is sharing the information because it believes it is important to be transparent with the public and the safety and security communities about this potential shift in capabilities. The company also said it is taking action, including enacting stricter security controls and pausing internal activities involving Astra that do not meet the beefed-up guardrails. It added that it is working with relevant government agencies and select AI safety organizations to test the model’s capabilities.

OpenAI says it slowed Astra model development over security concerns | TechCrunch

View Original