RecapAI
AI News
Seguridad

OpenAI says its Astra model is the first to cross a 'Critical' cybersecurity risk threshold

OpenAI says its Astra model is the first in the company's lineup to cross the 'Critical' cybersecurity risk threshold defined in its own Preparedness Framework — the internal system OpenAI uses to gate how much access to give a model based on its most dangerous capabilities. In expert-led testing against a hardened browser and operating system, Astra reportedly discovered previously unknown vulnerabilities on its own and chained them into working exploit chains, including a full browser-compromise chain that escaped the sandbox and executed commands on the host machine.

Under OpenAI's framework, 'Critical' is the threshold reserved for a model that can identify and build functional zero-day exploits against hardened real-world systems without human help, or independently devise and execute a full cyberattack strategy from just a high-level goal. Rather than shipping Astra broadly, OpenAI temporarily paused its development to add stronger safeguards — chain-of-thought monitoring, jailbreak detection, and containment-escape evaluations — before resuming, and says the model's strongest cyber capabilities will only be available to a vetted group of organizations in a cybersecurity coalition called Daybreak.

This isn't a ChatGPT feature update, but it matters for how you think about AI capability generally: it's a concrete example of a model demonstrably crossing from 'good at security research' into territory the lab that built it considers dangerous enough to gate deliberately, rather than ship by default. Whichever side of the AI-safety debate you land on, it's a useful real-world data point rather than a hypothetical one.

Related tools on RecapAI:ChatGPT

Originally reported by CNBC

View original source