OpenAI Halts Frontier Model 'Astra' After Cyber Capability Hits 'Critical' Threshold
In one line: OpenAI has reportedly held back its new frontier model, "Astra," after judging its cyber capabilities to have reached the "Critical" threshold.
Key points
- Astra's cyber-related capabilities are said to have hit the top risk tier ("Critical") of OpenAI's internal safety scheme, the Preparedness Framework.
- According to the report, this marks the first time the framework's highest level has been triggered by a frontier model.
- OpenAI's stated policy is not to deploy a model once it reaches "Critical" until further mitigations are in place.
Why it matters
If confirmed, this would be the first public case of an AI developer applying its own brakes based on a safety threshold rather than shipping commercially. It signals that a frontier model's ability to aid cyberattacks may have risen to a level that actually activates policy, with potential ripple effects on industry self-regulation and government AI rules.
Read more
- OpenAI Pauses Astra at ‘Critical’ Cyber Threshold — forkast.news