OpenAI has provided a first glimpse at its newest model, which the AI developer says is capable of discovering and exploiting previously unknown cybersecurity vulnerabilities without human intervention.
The firm said Astra, the foundation model for its next family of large language model (LLM) releases, is significantly more capable at cybersecurity tasks than its current flagship GPT-5.6 Sol, as well as more token efficient.
OpenAI believes Astra is the first model to meet the definition of “critical cyber capabilities”. The firm’s Preparedness Framework defines this as a model that can identify and develop working zero-day exploits in real-world environments without human intervention or plan and carry out end-to-end cyber attacks against “hardened targets” based only on a high-level end goal.
In a test run using the benchmark ExploitBench, which tests a coding agent’s ability to identify and exploit vulnerabilities, OpenAI said Astra scored 100 per cent compared to GPT-5.6 Sol’s 73.5 per cent and Claude Mythos’ 78 per cent.
The firm said it was unsure if this score reflected contamination, when models have been inadvertently trained on benchmark answers, so built an internal version of ExploitBench containing 20 recently discovered, high-severity vulnerabilities.
OpenAI said Astra achieved 39 per cent arbitrary code execution rates in this new benchmark, using 76,000 tokens, discovering and exploiting two previously undiscovered zero-days in the process. In comparison, GPT-5.6 Sol scored 11.5 per cent using 138,000 tokens.
AI models capable of advanced cyber attacks have dominated headlines in recent months and prompted government intervention. In July, OpenAI reported an “unprecedented cyber incident” which saw agents powered by GPT-5.6 Sol and an unreleased cyber model compromise the AI platform Hugging Face.
Subsequent reports from rival firm Anthropic revealed its own cyber-capable model, Claude Mythos, had carried out autonomous social engineering attacks.
OpenAI said it has put robust safety checks in place to prevent Astra from being misused or taking unauthorised malicious actions without human intervention. The firm has built these into its system stack and adjusted how it trains models, which it said has resulted in Astra refusing 91.5 per cent of cyber jailbreak requests compared to 59 per cent from GPT-5.6 Sol.
It also ran “honeypot” tests to check if Astra would repeat the Hugging Face incident by hacking security infrastructure to seek answers for a benchmark rather than completing it legitimately. Ultimately, it said, Astra showed no such behaviour.
OpenAI plans to provide Astra to alpha testers first, and will then make it accessible to trusted partners within its Daybreak Blue cybersecurity programme. This is OpenAI’s equivalent of Anthropic’s Project Glasswing.







Recent Stories