Home » OpenAI confirms Astra has reached ‘critical’ cyber threat level

OpenAI confirms Astra has reached ‘critical’ cyber threat level

by Brandon Duncan


OpenAI confirmed on Tuesday that its unreleased Astra model has reached a dangerous new milestone, while simultaneously confirming that it was forging ahead with a public launch.

In a blog post, OpenAI said that Astra has reached a “critical” cyber capability threshold, meaning the model could pose existential risks to cybersecurity. OpenAI’s Preparedness Framework tracks risk levels in three categories: biological/chemical, cybersecurity, and AI self-improvement.

OpenAI said that this is the first time any of its models has been evaluated at the critical level in the cyber domain.

The same blog post also stated that Astra will be “available soon,” but that its most advanced cybersecurity skills will be reserved for select testing partners, in the interest of public safety. The AI company said it was still preparing to safely release Astra and would be transparent about the potential threat level.

On X, Sam Altman addressed the tension between declaring Astra uniquely dangerous while also pushing ahead with a public launch. Altman said that “Astra has been done training for a while now” but that OpenAI has been “slowing things as needed to ensure that we can do sufficient work on safety and alignment.”

“There is an obvious tension here: on one hand, Astra is very good and we are excited to see what people will build with it,” Altman wrote. “On the other hand, we are clearly in a phase of development where we believe caution is warranted, and we are pacing our progress to ensure that we can meet the safety standards required by new capability levels.”

Of course, this isn’t the first time an AI company has hyped up one of its models as dangerously powerful ahead of a big launch.

“Limiting access to the most advanced cybersecurity features to select partners is a fair mitigation, and I don’t think OpenAI is being reckless here,” said Tal Kollender, Founder and CEO of AI cybersecurity firm Remedio. “Their framework is built to allow release with the right safeguards, but the security story is that defense hasn’t caught up to any version of this, gated or public.”

What makes Astra potentially dangerous?

OpenAI previously rated GPT-5.6-Sol as a “high” risk in the cyber domain, but Astra is even more capable. OpenAI says it scored 100 percent on the ExploitBench benchmarking test.

“Under our Preparedness Framework, a model reaches the Critical cybersecurity threshold if it can identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention, or can devise and execute end-to-end novel strategies for cyberattacks against hardened targets given only a high level desired goal,” an Aug. 7 OpenAI blog post stated.

In recent months, advanced frontier models from Anthropic and OpenAI have developed rapidly at agentic coding and cybersecurity hacking. As a result, the prospect of AI agent swarms hacking critical infrastructure no longer seems far-fetched, especially after the Hugging Face hack. In that incident, swarms of AI agents developed by OpenAI escaped a secure testing environment and hacked Hugging Face, acting autonomously in order to pass a test. Meanwhile, thanks to a deluge of AI-discovered bugs, some zero-day bug bounty programs have been forced to shut down entirely.

“While Astra was not involved in the Hugging Face incident, we have incorporated our learnings from that incident into our safety approach,” OpenAI’s blog post states. “Based on retrospective testing, we believe our production safeguards at the time would have prevented the Hugging Face incident. We have since implemented even stronger safeguards for Astra, including training the model to more reliably refuse harmful cyber requests and respect safety restrictions, additional protections against misuse, and monitoring that can stop potentially unauthorized activity.”

In its blog post, OpenAI also detailed some of the other safety precautions developed around Astra, with two goals in mind: preventing bad actors from accessing the model and stopping Astra from taking unwanted actions on its own. The company said it’s tightened its secure sandboxes and stepped up “offline detection and threat disruption” efforts.

“Nation-states and well-funded attackers aren’t waiting on Sam Altman’s release calendar,” Kollender told Mashable. “If a frontier lab’s internal model can find and chain zero-days without a human in the loop, assume adversaries are within a generation of the same capability, gatekept or not.”

On the same day OpenAI made these announcements, Anthropic announced the launch of Fable 5.1, an update to its latest frontier-level model. Fable is based on Claude Mythos, the model Anthropic deemed too dangerous to release because of its cybersecurity coding abilities.

If the pace of AI development is starting to remind you of War Games, keep one other fact in mind: While advanced frontier models do pose escalating cybersecurity risks, the same models will also benefit cybersecurity defenders in the long run.

UPDATE: Sep. 2, 2026, 1:26 p.m. EDT This article has been updated with comments from Sam Altman shared to X and quotes from a cybersecurity expert.

UPDATE: Sep. 2, 2026, 9:29 a.m. EDT A previous version of this article stated that Astra was the first OpenAI model to be evaluated as a “critical” threat in any of the three categories in the company’s Preparedness Framework (chemical/biological, cyber, self-improvement). The company has only said that Astra is the first model to reach the “critical” threshold in the cyber domain.


Disclosure: Ziff Davis, Mashable’s parent company, in April 2025 filed a lawsuit against OpenAI, alleging it infringed Ziff Davis copyrights in training and operating its AI systems.





Source link

You may also like

Leave a Comment