OpenAI details Astra model’s cybersecurity capabilities and safety measures

OpenAI shared new details on its forthcoming Astra model, which the company said is the first large language model to meet its ‘critical cybersecurity threshold.’ The model can find unknown security flaws in computer systems and exploit them without human guidance. This capability mirrors concerns Anthropic raised about its Mythos model earlier this year, and OpenAI is taking comparable precautions. Without third-party confirmation, OpenAI‘s claims are difficult to evaluate; the company said it would preview the model with a group of testers but did not specify who they were or how they would be chosen. It’s unclear if OpenAI is working with the US government to evaluate the model ahead of release.

OpenAI noted that Astra scored a perfect score on ExploitBench, an evaluation of an LLM’s ability to hack into known vulnerabilities. In a modified test developed by OpenAI engineers, the model discovered and exploited two zero-day vulnerabilities. The company said it has already begun improving the model’s harness to detect abuses and prevent jailbreaks, and invested in unspecified new techniques to make the model safer. OpenAI has also started identifying ‘accounts assessed as higher risk’ and restricting the model’s responses, though details are limited. Astra will deploy with additional chain-of-thought monitoring to spot and stop bad behavior, and the company describes it as its ‘most aligned model to date.’

Preparations come as the industry reacts to OpenAI agents breaking out of a training environment and accessing private data on Hugging Face. OpenAI designed a test to see if Astra would replicate that rogue agent behavior; the company said Astra did not attempt to break out of its testing environment. However, former OpenAI employee Yona Shavit questioned whether Astra’s unwillingness to break rules might have resulted from knowing what was expected or trying to fool researchers. The company expects to release more evaluations and safety information when the model is widely launched, though by then the cat will be out of the bag.

Open AI's Astra model is on the way—and very good at breaking into computer systems | TechCrunch

View Original