Another day, another major frontier AI release, except this particular OpenAI model is causing more jitters than usual in the AI community.
GPT-6 Astra, which OpenAI describes as its most advanced model to date, was announced Thursday afternoon, and for now it’s available only to members of OpenAI’s Daybreak cybersecurity program. The model will be released in the coming days to paid OpenAI and ChatGPT subscribers, including those on Pro, Plus, Enterprise, and Business accounts, as well as via the OpenAI API.
Arrived just days after the launch of Anthropic’s Claude Fable 5.1 and Mythos 5.1, Astra marks a “jump” in AI capabilities, OpenAI president Greg Brockman said, boasting that the new model “can really do anything a human can do with a computer.”
Astra is also OpenAI’s first model to reach the “critical” threshold of the company’s “preparedness framework” due to its extreme cybersecurity skills, meaning it could carry out “end-to-end” attacks on “hardened targets” on its own, among other capabilities.
OpenAI previously paused work on Astra to bolster its safeguards before announcing earlier this week that the model is “consistently more likely to respect explicit safety restrictions and warnings” than GPT-5.6 Sol, the OpenAI model involved in the now infamous Hugging Face attack.
Despite OpenAI’s assurances, AI experts remain worried about Astra. The new model is said to employ a reasoning technique known variously as “recurrent depth” or “opaque recurrence,” which (as TechCrunch describes) makes its “chain of thought” much harder to read.
Keeping tabs on a frontier AI model’s thinking is, obviously, a big deal when it comes to preventing the kinds of rogue AI hacks we’ve been hearing about over the past several weeks, and the potential of losing that kind of surveillance has spooked top AI researchers.
“If this is true, OpenAI seems to be violating one of the few redlines that exist in the AI community,” wrote Steven Adler, a former OpenAI safety lead, on X.
Buck Shlegeris, CEO of Redwood Research, echoed Adler’s worries. “I don’t know whether Astra is much less CoT [chain of thought] monitorable than previous models,” Shlegeris posted on X. “But if OpenAI pushes this technique further, they’ll have the option to massively increase the recurrence and totally destroy CoT monitorability.”
OpenAI chief scientist Jakub Pachocki has pushed back on the concerns. “OpenAI has worked to preserve and utilize chain-of-thought monitoring since our very first reasoning models,” Pachocki wrote. “We deeply care about this technique, as it can give us a view into how model alignment generalizes from its training distribution.”
Another former OpenAI researcher, Daniel Kokotajlo, responded to Pachocki that even if OpenAI doesn’t “go further” with recurrent depth reasoning, “others might.”
On Thursday, Pachocki suggested it wasn’t recurrent depth per se that was making new frontier AI models harder to monitor, but rather the fact that “more capable models can perform harder tasks using fewer language tokens” or even “no language tokens.”
I’ll be keeping an eye out for Astra to hit my ChatGPT plan, so stay tuned for my first impressions.



