Less than a week after touting the scientific achievements of Astra, its next “major” model, OpenAI says it’s “pausing internal activities” related to the model due to its powerful cybersecurity abilities.
“Our latest internal evaluations of Astra, one of our upcoming models, over the past few days indicate significant advancements in agentic coding and cybersecurity,” OpenAI stated in a Friday press release.
“These results, in addition to expert assessments, have led us to conclude last night that we cannot rule out critical cyber capabilities under our Preparedness Framework.”
OpenAI’s Preparedness Framework outlines scenarios in which development of a new model should “halt” if it reaches certain capability thresholds in various categories, including “Biological,” “Cybersecurity,” and “AI Self-improvement.”
For cybersecurity, the “critical” threshold means a model can pinpoint “zero-day exploits of all severity levels” in “hardened real-world systems” without any human help.
A model could also hit the “critical” level if it can carry out “end-to-end novel strategies for cyberattacks against hardened targets” with little more than a “high-level desired goal” in mind, according to the OpenAI safety framework.
OpenAI’s previous high-end model, GPT-5.6 Sol, only reached the “high” threshold during internal evaluations, the company said. OpenAI initially released GPT-5.6 Sol to just a “select group of trusted partners” before making the model public a couple of weeks later.
Given its concerns over Astra’s potential cybersecurity risks, OpenAI says is it “implementing stricter security controls” for the model, such as setting up “isolated testing environments” and “restricted network and tool access,” among other measures.
In the meantime, OpenAI is “pausing internal activities involving Astra that do not yet meet these strengthened security control requirements,” the company said.
OpenAI said it sounded its warning about Astra because “it’s important to be transparent to the public” about what Astra is potentially capable of.
Barely a week ago, OpenAI touted Astra’s abilities in mathematical research, including its solutions to 10 open math and computer science problems.
OpenAI’s revelations about Astra come amid a flurry of reports of advanced AI models going rogue, hacking real companies and organizations during training exercises and even forging phony credentials to hack external systems.
Now with Astra said to be demonstrating dangerously strong cybersecurity capabilities, it seems we may have reached a crossroads in AI safety, where each new “frontier” model on the AI test bench is judged — at least initially — too powerful to be released.



