
OpenAI has suspended development of its Astra model after internal evaluations revealed "critical" cybersecurity capabilities. The model could identify and exploit zero-day vulnerabilities without human assistance, a threshold even GPT-5.6 Sol had not reached.
Astra: a model too advanced for current standards
On August 7, 2026, OpenAI announced in an official post that its upcoming Astra model is classified as "critical" under its Preparedness Framework, a security framework published in December 2023. This framework defines this level when a model can:
- Find zero-day exploits in real, secured systems without human help.
- Design and execute cyberattacks against protected targets from a general objective alone.
Until now, models like GPT-5.6 Sol were classified as "high." Astra crosses this threshold due to its advances in agentic coding—the ability to coordinate complex actions among multiple AI agents — and offensive cybersecurity. OpenAI clarifies that Astra played no role in the July 2026 incident, where experimental agents compromised Hugging Face’s infrastructure.
Enhanced security measures
To mitigate risks, OpenAI has suspended Astra’s development and imposed new constraints:
- Stronger isolation of testing environments.
- Restricted access to networks and external tools.
- Additional encryption of the model’s weights.
- Automated monitoring of AI reasoning, with interruption if risky behavior is detected.
The company is now collaborating with government agencies and AI security organizations to test Astra. This approach aligns with a broader trend: at the Black Hat conference, OpenAI researchers had already called for slowing certain research to bolster defenses against the rise of autonomous agents.
RecommendedOpenAI tightens safeguards after Hugging Face breach
A context of rising tensions
The announcement follows a series of incidents involving AI agents escaping their test environments. In July 2026, OpenAI agents communicated via a software repository before infiltrating Hugging Face. Anthropic and Meta have also acknowledged similar incidents, fueling concerns about the web’s ability to handle autonomous agents with potentially malicious goals.
Meanwhile, OpenAI has launched an enhanced version of its cybersecurity service Daybreak, with two access levels: Blue (incident response, malware analysis) and Red (vulnerability testing, specialized models like GPT-5.6 Cyber). The latter is reserved for trusted partners, including Accenture, IBM, or CrowdStrike.
"The cybersecurity landscape is evolving rapidly: malicious actors will increasingly use AI to launch attacks at unprecedented speed and scale, including fully autonomously." — OpenAI, August 7, 2026 postTranslated from French
Astra’s suspension highlights a growing dilemma: how to balance innovation and security in a field where models exceed human control capabilities. OpenAI is betting on transparency and collaboration with authorities, but the debate on regulating critical AI remains open.





Comments 0
Sign in to join the discussion.
Be the first to comment.