OpenAI Will Limit Astra’s Cyber Tools After Rating the Model ‘Critical’ Under Its Own Safety Rules

Table of Contents

OpenAI plans to release its next model, Astra, soon, and to keep the model’s strongest cybersecurity functions off the open product at launch. The company said Tuesday that Astra is the first system it has formally rated at the Critical cybersecurity threshold in its Preparedness Framework.

That rating means that with tools and access, OpenAI believes Astra can find previously unknown flaws and build exploits against many well-protected real-world systems without a person directing each step. Amelia Glaese, OpenAI’s VP of research, used that wording with reporters on Tuesday.

The company paused parts of Astra’s development in early August after preliminary tests suggested the model might already sit at that level. It isolated test environments, restricted network and tool access, and stopped internal Astra work that did not meet the tighter rules. Astra was not the model involved in the summer Hugging Face incident. OpenAI still delayed Astra while it built refusals, jailbreak resistance, and monitoring meant to stop both a malicious user and an agent acting on its own.

What OpenAI says the tests showed

Astra beat GPT-5.6 Sol on vulnerability work and used fewer tokens to do it. On ExploitBench, which asks a model to build exploits from known bugs, OpenAI reported a perfect score. On an internal set of 20 high-severity V8 JavaScript engine flaws disclosed between June and August 2026, testers said Astra found and used two zero-days in a chain. OpenAI said it is disclosing those two bugs to the maintainers.

Under the 2023 Preparedness Framework, Critical is the top cyber rung. One path to it is identifying and developing working zero-days of all severity levels across many hardened systems without human intervention. OpenAI now says Astra meets that bar. It also says the new safeguards “sufficiently minimize the risk of severe harm for release.” It did not give a ship date.

Two products, not one

Most users will get a version with stronger refusal and monitoring. If the system flags possible cyber misuse or an unauthorized action, ChatGPT and Codex users may have to approve the next step. API tasks stop. OpenAI warned that the same filters can interrupt work that has nothing to do with hacking, including long-running agent jobs.

The less restricted cyber stack goes first to a small tester group: people and organizations that protect digital and other critical infrastructure, including the U.S. government and firms already in OpenAI’s trusted cyber access program. OpenAI declined to name them. Daybreak Blue, the company’s controlled-access defensive program, is the next expansion. Wired reported that Daybreak partners have included Cisco, Cloudflare, and Palo Alto Networks.

OpenAI is selling defensive cyber use as a line of business. Restricting offense-shaped skills to vetted defenders is how it tries to keep that sale from becoming a general exploit factory.

What security and privacy teams should do with this

If your company uses ChatGPT, Codex, or the API, assume Astra-era monitoring will kill or pause some agent runs. Write that into runbooks before a Friday batch job dies mid-flight.

If you want the defensive capabilities, Daybreak Blue and the alpha tester list are the doors. Expect contractual limits on who may prompt for exploit work and logging of those sessions.

If you maintain software Astra was tested against, watch for the two zero-day disclosures OpenAI said it is sending. Treat “a frontier lab’s eval set” as a reason to patch faster, not as a press mention.

OpenAI’s own rulebook forced the delay. The company is now trying to ship the model anyway by splitting access. Whether refusals and a partner list hold once the weights and tools are in more hands is the part that will not be settled by a blog post.

Online Privacy Compliance Made Easy

Captain Compliance makes it easy to develop, oversee, and expand your privacy program. Book a demo or start a trial now.