NC / HOME NEWS

OpenAI Says Astra Crosses Its Critical Cyber Threshold

OpenAI says its upcoming Astra model can find and exploit unknown flaws, so advanced cyber capabilities will launch behind tighter controls.

Representative data-center server racks illustrating infrastructure around OpenAI’s Astra cybersecurity model
Image: BalticServers.com via Wikimedia Commons, CC BY-SA 3.0

OpenAI says its forthcoming Astra model has crossed the company’s Critical cybersecurity capability threshold, a designation that changes how the model must be developed and released. In its September 1 safety update, OpenAI says Astra can find previously unknown security flaws and develop exploit chains across well-protected systems with the right tools and access, without a person guiding each step.

Astra is not publicly available yet. OpenAI says it plans to release the model soon, but its most advanced cybersecurity work will initially be limited to a small group of testers, with access through the Daybreak Blue program expanding afterward. CNBC’s report likewise describes the announcement as a limited rollout rather than a general launch.

What OpenAI’s threshold means

Under OpenAI’s Preparedness Framework, the Critical threshold covers models that can identify and develop functional zero-day exploits across many hardened real-world critical systems without human intervention. It also covers models that can devise and execute novel, end-to-end cyberattack strategies against hardened targets from a high-level goal. OpenAI says Astra met the threshold through a mix of public and private benchmarks plus expert-led assessments.

One published example is ExploitBench, where OpenAI says Astra scored 100% on a benchmark measuring exploit development from known vulnerabilities. To address contamination concerns, the company also created an internal test with 20 more recently disclosed, high-severity vulnerabilities. OpenAI says Astra achieved higher arbitrary code-execution rates than GPT-5.6 Sol on that test while using fewer output tokens, and that it discovered and used two zero-day vulnerabilities as part of an exploit chain. The company says those vulnerabilities are being disclosed to maintainers.

OpenAI also describes expert assessments involving a hardened browser and operating system. According to the company, Astra built a browser-compromise chain that escaped a sandbox and executed commands on the host, then combined operating-system vulnerabilities into a local privilege-escalation chain. These are OpenAI’s internal evaluation results, not an independent audit, and the company says Astra’s results reflect Daybreak Blue access rather than the default production configuration.

The release comes with more controls

OpenAI says it delayed parts of Astra’s development and release while strengthening isolation, network controls, model-weight protection, monitoring, and alignment training. The work followed the company’s response to the Hugging Face incident, although OpenAI says Astra was not involved in that incident. It also says a large frontier reinforcement-learning run restarted on August 28 after the new requirements were in place, while some smaller experimental runs remain paused.

At the model layer, OpenAI reports that Astra refused 91.5% of requests in its cyber-jailbreak evaluations, compared with 59% for GPT-5.6 Sol. It says higher-risk accounts will face a more conservative behavior boundary, and that production monitoring will inspect reasoning and actions for unauthorized behavior. If the monitor flags a task, ChatGPT or Codex users may be asked to review it, while API tasks can stop instead.

The news extends our earlier coverage of OpenAI’s GPT-5.6 family and its model safeguards, while the company’s text watermarking report showed how model deployment questions reach beyond raw capability. Astra adds a sharper security test: the system can be useful to defenders while also creating a higher consequence if access controls fail.

OpenAI says it will publish more detail in Astra’s system card at launch. Until then, the important boundary is clear: Astra is an announced, not yet generally available model, and the capability claims come from OpenAI’s evaluations. The limited tester plan and additional monitoring are part of the release design, not evidence that the model is already available for ordinary users.