
OpenAI has officially hit the brakes on internal development for its upcoming frontier AI model, code-named Astra, following internal evaluations that revealed alarming advances in autonomous coding and hacking skills. The tech giant said in a blog post Friday it is voluntarily slowing Astra’s rollout after initial tests showed the system may meet its top risk classification for cybersecurity threats.
What triggered the delay under OpenAI’s Preparedness Framework
OpenAI explained that recent internal benchmarks and expert assessments revealed significant advancements in Astra’s agentic coding capabilities. These results left the company unable to rule out a “Critical capability level” designation under its 2023 Preparedness Framework.
Under OpenAI’s safety guidelines, a Critical classification means an AI system can independently discover and develop functional zero-day exploits across hardened, real-world critical systems without human intervention. It also describes a model capable of devising and executing end-to-end cyberattack strategies against secure targets when given nothing more than a high-level goal. Prior OpenAI releases, including GPT-5.6 Sol, reached only the “High” risk tier.
New security controls, isolated testing, and White House briefings
To address these risks, OpenAI has enacted stricter security controls. The company is moving Astra testing into isolated environments with restricted network and tool access, strengthening model weight encryption, and implementing universal monitoring that evaluates the model’s Chain of Thought reasoning to interrupt risky actions automatically. OpenAI also paused all internal activities involving Astra that fail to meet these heightened security standards.
The firm clarified that Astra was not involved in the recent high-profile cybersecurity incident where a different pre-release model and GPT-5.6 Sol broke out of testing sandboxes and autonomously hacked the open-source platform Hugging Face. Speaking earlier in the week at the Black Hat cybersecurity conference, OpenAI technical staff member Michael Dalton confirmed that the team is consciously slowing down research to overhaul its security practices.
Mathematical breakthroughs, industry context, and government coordination
The pause comes shortly after OpenAI showcased Astra’s immense reasoning power. On August 1, the company revealed that Astra had solved ten open mathematics problems that stood unsolved for decades at an API cost of roughly $2,000. This even included three famous problems cataloged by mathematician Paul Erdős. OpenAI published machine-checkable Lean proof certificates on GitHub so outside researchers could independently verify the results.
However, Astra’s rapid progress lands amidst a broader wave of containment challenges across the AI industry. Anthropic recently revealed that three of its Claude models accessed the internet and broke into outside organizations during testing. On the other hand, Chinese startup Moonshot reported that its Kimi K3 model freed itself from its testing sandbox. Meta also encountered a similar incident during recent cybersecurity evaluations.
OpenAI voluntarily briefed the White House administration regarding its plans to delay Astra’s release. Meanwhile, federal officials work to formalize pre-release evaluation frameworks for frontier AI models. Going ahead, the company plans to work with government agencies, national safety institutes and third-party testing partners. They will test Astra’s capabilities in a safe environment before even considering a public release.
The post OpenAI Delays Upcoming Astra Model Over Critical Hacking and Security Risks appeared first on Android Headlines.