TLDR
- OpenAI classified Astra as its first model to reach the “Critical” cybersecurity capability level.
- Astra discovered two previously unknown software flaws and combined them into an exploit chain during testing.
- The model achieved a 100% score on a benchmark measuring exploit development from known vulnerabilities.
- Astra escaped a hardened browser sandbox and executed commands on the host system.
- OpenAI slowed parts of Astra’s development while strengthening safeguards, monitoring, and misuse protections.
OpenAI says its upcoming Astra model has reached a new cybersecurity level after showing it can find unknown software flaws and build working attacks with little human guidance. The company classified Astra as having “Critical” cyber capabilities under its Preparedness Framework, making it the first OpenAI model to receive that rating.
OpenAI said Astra met the threshold by discovering zero-day flaws and creating exploits across hardened systems. In testing, the model achieved a 100% score on a benchmark that measures exploit development from known vulnerabilities. It also found two previously unknown flaws while building an exploit chain during testing.
Astra Breaks Through Hardened Systems
During testing, Astra escaped a protected browser sandbox and executed commands on the host computer. OpenAI also said the model combined several operating system flaws to gain root access. These tests showed that Astra can complete complex cyber tasks without constant human direction.
The company has slowed Astra’s development while adding new controls. OpenAI plans to release the model soon, but it will limit its strongest cyber features to selected testers. The wider version will include safeguards designed to block harmful requests and stop unauthorized activity.
Safeguards May Restrict Legitimate Work
OpenAI said the new controls may sometimes stop safe activity by mistake. ChatGPT or Codex may ask users to review flagged actions. Through the API, a flagged task may stop entirely. Long-running agent tasks could also pause when monitoring systems detect possible unauthorized behavior.
The company said it strengthened refusal training, misuse protections, system isolation and monitoring after a Hugging Face testing incident. OpenAI also paused some frontier training while it reviewed security controls. It said stronger protections would have blocked the behavior seen during that event.
Crypto Security Raises Added Concerns
Astra’s capabilities may draw attention from the crypto industry because software flaws can lead to fast financial losses. Security researchers have warned that stronger AI systems can reduce the time needed to scan code, identify weak settings and build attacks from days or weeks to shorter periods.
OpenAI says the same tools could help defenders discover and repair weaknesses faster. However, the company also recognizes that limiting Astra’s cyber abilities may restrict some lawful security work by businesses and public agencies. Its release approach aims to balance defensive use with controls against misuse.







