OpenAI Designates Unreleased ‘Astra’ Model as First With ‘Critical’ Hacking Capabilities

A robotic hand reaching toward the OpenAI logo on a white background

Quick Read

  • Astra is the first OpenAI model to reach the 'critical' cybersecurity threshold.
  • The model can independently create zero-day exploits and execute complex cyberattacks.
  • Astra achieved a 100% pass rate on the ExploitBench testing framework.
  • Access to the model's advanced tools is gated for alpha testers and defensive security partners.
  • OpenAI briefly paused Astra's development in August 2026 due to rapid capability growth.

A New Threshold in AI Capability

OpenAI announced Tuesday that its unreleased AI model, Astra—widely speculated to be the foundation for GPT-6—has officially crossed the “critical” cybersecurity capability threshold under the company’s internal Preparedness Framework. This designation marks the first time any OpenAI model has been classified at this level, signaling a significant shift in the autonomous potential of large language models (LLMs).

According to OpenAI, the “critical” classification is reserved for systems that can independently develop functional zero-day exploits across multiple hardened real-world systems or plan and execute entire cyberattacks starting from high-level objectives. Previous iterations, including the current top-tier model GPT-5.6 Sol, were capped at the “high” tier of the framework.

This report draws on information published by tech.yahoo.com.

Performance and Benchmarking

The model demonstrated significant technical proficiency during internal evaluations. In testing against ExploitBench, a specialized framework designed to measure the ability to turn known software vulnerabilities into functioning exploits, Astra achieved a perfect 100% success rate. To ensure these results were not the product of memorization, researchers conducted additional testing using 20 high-severity vulnerabilities identified in Google’s V8 JavaScript engine between June and August 2026. Astra outperformed GPT-5.6 Sol while consuming fewer computational resources, successfully chaining two previously unknown zero-day vulnerabilities during the process.

In simulated environments, Astra demonstrated the ability to break out of browser sandboxes and execute arbitrary code on host operating systems after merely opening a malicious HTML file. It further successfully identified and chained multiple vulnerabilities to achieve privilege escalation from a standard user account to root-level access.

Safeguards and Deployment

Despite these capabilities, OpenAI reports that Astra has shown improved resilience against adversarial manipulation, refusing approximately 91.5% of cyber jailbreak attempts in internal testing, compared to 59% for GPT-5.6 Sol. The company has announced that access to these advanced features will be strictly gated. A small group of alpha testers will receive initial access, followed by a broader rollout through the “Daybreak Blue” program, which is specifically tailored for defensive security research.

The announcement follows a period of heightened internal caution. OpenAI had briefly paused Astra’s development in August 2026 following rapid advancements in its coding and security skills. While the industry has been monitoring competition from rivals like Anthropic—which released its own Fable 5.1 and Mythos 5.1 models on September 1—OpenAI has yet to confirm a public release date for Astra, though prediction markets have shifted expectations toward late 2026.

|
Contributor:Azat TV Editorial
|
Publisher:Azat TV

LATEST NEWS