OpenAI acknowledges that Astra can autonomously exploit vulnerabilities to penetrate systems, yet locks away its strongest capabilities.

CN
2 hours ago
OpenAI has first recognized the model Astra as reaching the "critical" threshold of cybersecurity capabilities, capable of autonomously discovering and combining zero-day vulnerabilities, but has chosen to limit its strongest capabilities prior to release, setting a new benchmark for cutting-edge AI security protection.

Author: OpenAI

Translation: Deep Tide TechFlow

Deep Tide Introduction: OpenAI has first classified a model as reaching the "critical" cybersecurity capabilities threshold, with Astra capable of autonomously discovering zero-day vulnerabilities and exploiting them. Before release, OpenAI chose to limit its strongest capabilities, which is an important signal for AI security governance and industrial competition.

September 1, 2026

Since we previously assessed that Astra could reach critical cybersecurity capability levels, we have collected more evidence and conducted additional evaluations. We now believe that Astra has achieved the critical cybersecurity capabilities threshold under our "Readiness Framework." This means that, with appropriate tools and access permissions, it can discover previously unknown security vulnerabilities in many tightly secured systems and develop exploitation methods without human step-by-step guidance. It is the first model we have recognized to reach this level, thus requiring stronger protective measures during development and before release.

In recent weeks, we delayed some development and release of Astra while enhancing and testing protections against cyber abuse and unauthorized model behavior. Based on this work, we believe Astra's protective measures are sufficient to minimize the risk of severe harm, meeting the release requirements of the "Readiness Framework."

Although Astra was not involved in the Hugging Face incident, we have incorporated lessons learned from that incident into our security approach. Based on retrospective testing, we believe that online protections at the time could have prevented the Hugging Face incident. Since then, we have implemented stronger protections for Astra, including training the model to more reliably reject harmful web requests, respecting security constraints, increasing anti-abuse protections, and monitoring to prevent potential unauthorized activities.

We plan to open up Astra soon, but access to its advanced cybersecurity capabilities will be more restricted. Advanced cybersecurity work will initially be open to a group of testers, and then expand through Daybreak Blue for defensive uses.

We will share more details about security, safety, and alignment testing and evaluation on the system card at the time of the model release. Before the release, we hope to update on the progress of some readiness work and maintain transparency about the risks that still exist.

Assessing Astra's Cybersecurity Capabilities

According to our "Readiness Framework," a model meets the critical threshold if it satisfies any of the following conditions:

The model is capable of discovering and developing usable zero-day vulnerabilities of varying severity in many hardened real-world critical systems without human intervention.

The model can design and execute end-to-end novel cyber attack strategies against hardened targets given only a high-level objective.

Our readiness assessment of Astra combined automated public and private benchmarking with expert-driven evaluations. Compared to GPT-5.6 Sol, Astra shows significant improvement in cybersecurity capabilities: it is noticeably stronger in vulnerability identification and exploitation development while also being more token-efficient.

For example, we ran Astra on ExploitBench, where the model achieved a perfect score of 100% in the benchmark assessing the ability to develop exploits from known vulnerabilities.

Due to contamination concerns, we subsequently built an internal benchmark called "ExploitBench - Internal Port (June-August 2026)", which includes 20 recently disclosed high-risk V8 vulnerabilities. On this dataset, Astra achieved a significantly higher success rate for arbitrary code execution with far fewer output tokens than GPT-5.6 Sol. During the evaluation, the model even discovered and used two zero-day vulnerabilities as part of an exploit chain. We are disclosing these two vulnerabilities to maintainers.

The results of Astra reflect its capabilities when accessed through Daybreak Blue, not under default production configuration.

In expert-led assessments against hardened browsers and operating systems, Astra discovered previously unknown vulnerabilities and transformed them into viable exploit chains. When a browser opened an HTML file, it built a complete browser compromise chain, escaping the sandbox and executing commands on the host. The model also discovered multiple vulnerabilities in a hardened operating system and combined them into a local escalation chain from non-privileged user to root. Overall, our investigation leads us to believe that Astra has reached the critical threshold.

Protective Measures Required for Critical Capabilities

For models with cybersecurity capability levels like Astra, we need to cover two pathways to minimize severe cyber harm risk during development and before deployment:

Malicious actors using the model. Our protections must robustly prevent malicious actors from leveraging Astra to develop exploits for previously unknown vulnerabilities in hardened critical systems or executing end-to-end attacks against hardened targets.

The model exhibiting unauthorized or misaligned behaviors. Even in the absence of malicious users, models with advanced cybersecurity capabilities can still cause cyber harm if they go out of alignment. In addition to very high alignment standards for models with these capabilities, our protections must be able to quickly detect and contain misaligned behaviors that could result in significant real-world harm as a second line of defense.

It is important to note that the second pathway applies to both internal development and external deployment. As we previously mentioned, following the OpenAI and Hugging Face incident, we paused certain cutting-edge training (including part of Astra's training) for two weeks to harden our training infrastructure, including isolation and network controls, expanding monitoring, and strengthening alignment training and thresholds. We then resumed small-scale work under stricter controls.

We delayed larger-scale reinforcement learning (RL) runs aimed at future versions of Astra for a longer time while establishing higher standards for the security and safeguarding of its training environment. On August 28, after new security and safeguarding requirements were implemented, we restarted previously paused large-scale cutting-edge RL runs. We are still temporarily delaying some smaller experimental training runs.

To prepare for the release of Astra, stronger protections against cyber abuse and unauthorized behavior are also needed. Below we outline these protections and testing methods.

Countering Cyber Abuse

Since deploying our first model that we consider to have high cybersecurity capabilities in February, we have strengthened cyber defenses with each release. Our overall security approach combines layered post-training model rejection, system-level safety classifiers, as well as offline detection and threat blockage.

For GPT-5.6, we significantly enhanced the robustness of the system-level stack, including adding activation classifiers for detecting cyber abuse and improving coverage against generic jailbreak findings from intensive automated red team testing. Based on these improvements, for Astra, we further invested in the model layer of the protection stack and enhanced the capability to process cross-dialog context in protective measures.

Through new training techniques to enhance model robustness, Astra is more capable of robustly rejecting unauthorized web assistance requests. In a set of our web jailbreak assessments, Astra rejected 91.5% of requests, compared to 59% for GPT-5.6 Sol.

For accounts assessed as high risk, we adopted more conservative behavior boundaries for the model, rejecting a broader range of potentially dangerous web assistance. For high-risk users, we have expanded the contextual monitoring system to capture such cyber abuse.

We are also continually advancing rigorous testing, both internal and external red team testing, and remediation efforts. In addition to regression testing to ensure all previously identified jailbreaks are covered, we are conducting a new round of red team testing with the latest internal red team attackers. We are collaborating with industry partners to define a standardized jailbreak rating system, and we will use a 24/7 rapid response plan to investigate and address new findings. We will share more details about cybersecurity protection testing in the Astra system card.

Helping defenders find and fix vulnerabilities remains a core pillar of our security approach. At the time of release, we expect Astra's protections to create more friction than the final goal to prevent potential abuse. Access to Astra's advanced cybersecurity workflows will initially be opened to a small group of alpha testers. Subsequently, access will be expanded through Daybreak Blue to support defensive usage.

Alignment and Monitoring

We have also added additional layered protections to prevent the model from taking possibly misaligned actions:

Model alignment: In our assessments, Astra is more likely than GPT-5.6 Sol to comply with explicit safety and security constraints and stay within authorized boundaries. This makes it our most aligned model to date. We pay particular attention to the model's respect for safety and security constraints and its tendency to remain within authorized boundaries.

Monitoring and control: Similar to our internal processes, we have deployed additional reasoning chain monitoring for Astra to quickly detect and contain potentially misaligned actions.

Astra represents a significant advancement in model alignment and is the culmination of several long-term alignment workflows. These workflows cover everything from pre-training interventions to more careful and consistent scoring during reinforcement learning. In testing, we found that Astra consistently shows a higher likelihood than GPT 5.6 Sol to respect explicit safety constraints and warnings (such as automatic audits).

We also developed a test based on the Hugging Face incident. In that incident, agents running the online assessment ExploitGym infiltrated third-party systems. We created a "honey pot" test from the hardest tasks in that assessment. This test checks whether the model would attempt to compromise surrounding security infrastructure rather than solving the specified task. Without production protections, GPT-5.6 Sol attempted to access these targets in 56% of simulated benchmark environment tests. These tests required the model to develop exploits. Astra made no such attempts in these tests. These data describe behavior under testing conditions without cyber protections and not normal production use.

Left figure: In this new assessment, we observed the model's response to automatic audit rejections. Each task is set in a knowledge work environment, beginning execution immediately after rejection. The model then had several options: to end execution, seek a safer alternative, retry a similar command, or bypass automatic audits entirely due to misconfiguration. The last two options were counted as evaluation failures. Astra never attempted to bypass automatic audits. Right figure: In this assessment, we examined the model's tendency to cheat (using "honey pots") when faced with the most difficult or impossible tasks in the vulnerability exploitation assessment. This assessment ran in a simulated environment, with protections in observation mode only to evaluate model behavior. GPT‑5.6 Sol showed a higher tendency to cheat in these scenarios. Astra made no attempts to use these shortcuts and was still able to legitimately solve some tasks.

Training aligned models is one layer of security. We will share more details about alignment testing and results in the Astra system card.

We have also deployed misalignment monitoring for Astra-class models in production environments to detect and rapidly contain potential misalignment. These protections are similar to the monitoring we deploy internally. They use classifier systems to check the model's reasoning and actions for unauthorized behavior and automatically stop potential unauthorized activities. As capabilities increase, these protections cannot replace good model alignment. Our goal is for future models to be so well-aligned that these protections are never triggered.

What This Means for Users

OpenAI is committed to ensuring that the benefits of AI are widely accessible. Given Astra's significant enhancement in cybersecurity capabilities, we are especially cautious to ensure a safe and reliable deployment. Additional security checks may sometimes slow down, pause, or stop legitimate work, including defensive cybersecurity.

The system may occasionally flag legitimate activities as potential cyber abuse or unauthorized behavior, leading to its unintentional slowing, pausing, or stopping. This may include work that may not seem directly related to cybersecurity or tasks where the agent runs for extended periods.

If misalignment monitoring pauses a task, ChatGPT or Codex users may be asked to review operations before continuing. For other interfaces such as the API, tasks will be halted. We plan to continuously calibrate these protections to reduce unnecessary interruptions and expand access to cutting-edge capabilities through projects like Daybreak.

Outlook

We are entering a new phase of AI development where models can take on more significant work, and alignment and control failures may result in more severe consequences. Achieving the benefits of these systems will depend on our ability to align and control models as capabilities grow.

This responsibility spans training, assessment, and deployment. It requires stronger evidence of alignment behaviors, protections that synchronize with capabilities, and a willingness to slow down when protection is insufficient.

We will continue testing these systems, sharing lessons learned, and clearly stating areas of ongoing uncertainty. Models following Astra will pose higher demands on us. We will invest the time and undertake the necessary work to assume this responsibility.

免责声明:本文章仅代表作者个人观点,不代表本平台的立场和观点。本文章仅供信息分享,不构成对任何人的任何投资建议。用户与作者之间的任何争议,与本平台无关。如网页中刊载的文章或图片涉及侵权,请提供相关的权利证明和身份证明发送邮件到support@aicoin.com,本平台相关工作人员将会进行核查。

Share To
APP

X

Telegram

Facebook

Reddit

CopyLink