OpenAI pauses Astra work after evals cannot rule out critical cyber capability
OpenAI said internal evaluations of Astra—an upcoming model not involved in the Hugging Face incident—show significant advances in agentic coding and cybersecurity, and that it cannot currently rule out Critical cyber capability under its Preparedness Framework (autonomous zero-days in hardened systems or end-to-end novel attacks from a high-level goal). The company is scaling safeguard and security-control testing, implementing stricter controls for higher-capability models (isolated eval environments, restricted network/tool access, enhanced weight protection, sandboxed execution), pausing internal Astra activities that do not yet meet those requirements, adding universal monitoring of Chain-of-Thought for risky/misaligned agentic actions, and planning government and AI-safety-org testing before any deployment.






