OpenAI Slows Astra Development After Critical Cybersecurity Review

OpenAI slowed development of its upcoming Astra model after internal tests identified advanced cyber capabilities. The company introduced stricter controls, paused activities that failed its security requirements, and began working with government agencies and safety organizations to assess the model.
OpenAI Slows Astra Development After Critical Cybersecurity Review
Written By:
Kelvin Munene
Reviewed By:
Achu Krishnan
Published on
Updated on

OpenAI has slowed work on its upcoming Astra model after internal tests raised cybersecurity concerns. The company said Astra may have reached a critical capability level under its safety rules.

The model displayed advanced agentic coding and cyber skills during preliminary testing. OpenAI said Astra could potentially identify vulnerabilities and carry out attacks against protected real-world systems without continuous human direction.

OpenAI Activates Critical Cyber Safeguards

OpenAI’s internal review placed Astra near the highest cyber-risk category in its Preparedness Framework. The company introduced that framework in 2023 to track emerging risks from advanced artificial intelligence models.

The framework requires stronger protections when a model approaches critical capability levels. OpenAI said, ‘We cannot rule out a critical capability level at this time.’ The company will continue testing Astra before making a final assessment.

OpenAI has paused internal Astra activities that do not meet its stricter security requirements. However, the company has not stopped all development or announced a new release date for the model.

The new controls include isolated testing environments and monitoring across Astra’s agentic applications. These measures aim to limit unauthorized actions during evaluations and prevent the model from operating outside approved systems.

Astra Release Faces Additional Testing

OpenAI plans to expand its safety testing before considering any public release. The company will assess how Astra identifies system vulnerabilities, writes attack code, and completes cyber tasks without direct assistance.

According to OpenAI, Astra remains under development and did not participate in the recent Hugging Face security incident. That case involved another unreleased model during an internal cybersecurity evaluation.

The clarification separates Astra’s testing from earlier cases involving models that moved beyond controlled environments. Those incidents increased questions about how AI laboratories manage autonomous cyber tools during development.

OpenAI said public disclosure allows researchers and security specialists to prepare for changing AI capabilities. The company also wants outside experts to review its testing methods and proposed safeguards.

Michael Dalton, a member of OpenAI’s technical staff, described the approach during the Black Hat cybersecurity conference. He said the company was “consciously slowing down research to enhance security” while it upgraded testing practices.

Government Agencies Join Astra Review

OpenAI has started working with relevant government agencies and selected AI safety organizations. These groups will help examine Astra’s cyber performance and review the protections surrounding its continued development.

A White House official said OpenAI voluntarily informed the administration about its plans to delay Astra’s release. However, OpenAI has not provided a public schedule for completing the additional evaluations.

The Trump administration is also developing a process for reviewing advanced AI models before release. Industry representatives received information about the proposed framework during recent briefings, although several operational details remain unresolved.

Questions include how developers will contact government reviewers and how long each assessment could take. Officials must also determine who can access unreleased models during national security evaluations.

Meanwhile, other AI developers have adjusted releases after detecting advanced cyber capabilities. Anthropic released a more restricted version of its cyber-focused Mythos model in June after applying additional safeguards.

OpenAI said it chose to disclose Astra’s status to maintain transparency with the public and security community. Development will proceed only under the stronger controls required for models approaching its critical cybersecurity threshold.

Also Read: Canada Opens AI Transparency Consultation Over Chatbots and Deepfakes

Join our WhatsApp Channel to get the latest news, exclusives and videos on WhatsApp
logo
Analytics Insight: Top Tech & Crypto Publication | Latest AI, Tech, Crypto News
www.analyticsinsight.net