OpenAI Designates Astra Model at Critical Cybersecurity Capability Threshold
OpenAI reports Astra achieved a perfect score of 100% on internal ExploitBench evaluations and refuses 91.5% of requests on cyber jailbreak evaluations. Traini…
Published on MyPrivateClaw
Sep 2, 2026, 4:20 PM UTC
Coverage date
Sep 1, 2026
Last updated
Sep 2, 2026, 4:20 PM UTC
News summary
OpenAI has designated its upcoming model Astra as meeting the Critical cybersecurity capability threshold under its Preparedness Framework, making it the first model OpenAI designates at this level. On the ExploitBench benchmark evaluating exploit development from known vulnerabilities, Astra achieved a perfect score of 100%. On an internal ExploitBench port containing 20 high severity V8 vulnerabilities disclosed more recently, Astra discovered and used two zero day vulnerabilities as part of an exploit chain during evaluation. OpenAI temporarily paused reinforcement learning training on Astra for two weeks following the OpenAI Hugging Face incident and restarted its largest planned frontier RL run on August 28, 2026 after implementing strengthened security controls. OpenAI plans to make Astra available soon, with advanced cybersecurity capabilities initially limited to a group of test…