Skip to content

UPDATED 20:58 EDT / AUGUST 07 2026

AI

OpenAI reveals upcoming Astra model may possess ‘critical’ hacking capabilities

OpenAI Group PBC today disclosed that one of its unreleased large language models may pose a significant cybersecurity risk.

The algorithm, which is known as Astra, was first detailed last week. OpenAI revealed in a Sunday blog post that the LLM had solved 10 long-running math problems. The company published the proofs and revealed that each one took about $2,000 worth of tokens to generate.

As part of its artificial intelligence safety efforts, OpenAI has published a 29-page document known as the Preparedness Framework⁠. One of the document’s sections contains a rating system for AI risks. The system ranks the cybersecurity risks posed by an LLM as “High” or “Critical” depending on its capabilities.

OpenAI’s flagship GPT-5.6 Sol model and a few earlier algorithms were given a High rating. According to the company, Astra is the first of its LLMs that may qualify for a Critical designation. Its engineers drew that conclusion based on a series of recent cybersecurity tests.

“These results, in addition to expert assessments, have led us to conclude last night that we cannot rule out critical cyber capabilities,” OpenAI stated in a blog post.

Under the Preparedness Framework⁠, a model poses a Critical cybersecurity risk if it can find zero-day exploits in “many hardened real-world critical systems.” The model qualifies if those exploits span multiple severity levels and are discovered without any human assistance.

OpenAI also designates an LLM as Critical if it can launch cyberattacks against hardened systems based on only a high-level hacking goal provided by a user. The company didn’t specify which of its two designation criteria were met by Astra. However, it did share details about how it’s tackling the risk.

The company is taking steps to ensure that Astra can’t access the public web. According to the company, its engineers will run the model in test environments that have restricted network and tool use permissions. OpenAI is pausing development activities that aren’t carried out in such sandboxes.

It’s also stepping up its efforts to prevent hackers from stealing Astra’s code. According to today’s blog post, the initiative will place particular emphasis on the encryption that protects the LLM’s weights. Those are the configuration settings that determine how a model processes data.

Astra powers a number of internal AI agents. The company has implemented observability mechanisms that monitor those agents for malicious activity. The mechanism spot suspicious behavior by analyzing agents’ chain of thought, a step-by-step summary of inference activities.

OpenAI will share some of the cybersecurity workflows it has developed with “third-party testing partners.” Those partners help the ChatGPT developer run the sandboxes in which it evaluates its LLMs’ capabilities. Additionally, OpenAI plans to loop in relevant government agencies and AI safety organizations. 

Image: OpenAI

Send us a news tip

Send us a News Tip

  • This field is for validation purposes and should be left unchanged.

Sign in or create an account

SIGN IN

OR

New User? SIGN UP

Join us

SIGN UP

Bio

Ethics statement

Extract the signal from the noise

Get SiliconANGLE updates and analysis.

Contact us

Partner with us

Contact us

Guest inquiry