OpenAI Slams the Brakes on Astra as Cyber Risks Outrun Its Safeguards
OpenAI Slams the Brakes on Astra as Cyber Risks Outrun Its Safeguards
OpenAI’s next major model is powerful enough to make its own developers hesitate. Astra may be headed for broad release, but first the company is slowing work to contain a cybersecurity risk it says it cannot yet dismiss.
The tension has been building since OpenAI published its Preparedness Framework in December 2023, meant to trigger tougher controls as frontier capabilities emerge. Previous systems, including GPT-5.6-Sol, were judged at the lower “High” cyber threshold; Astra’s recent internal tests marked a sharper turn. OpenAI said it “cannot rule out critical cyber capabilities” after seeing major gains in agentic coding and cybersecurity.
Under the company’s definition, that threshold includes independently finding and developing zero-day exploits across hardened critical systems, or carrying out novel end-to-end attacks from a high-level instruction. Reporting on the announcement framed the move as a potentially unusual step: a frontier lab slowing its own model over cyber concerns.
OpenAI stressed that Astra was not involved in the Hugging Face exploitation episode, which had intensified scrutiny of autonomous AI behavior. It has now paused Astra-related internal work that fails tougher requirements, while adding isolated test environments, restricted access, stronger protections for model weights and universal monitoring of agentic activity.
Greg Brockman described the evaluations as showing “significant capability advancements in agentic coding and cybersecurity,” while arguing the work is aimed at putting those capabilities “into the hands of defenders.” Sam Altman struck a similar balance: Astra should be generally available, he said, because keeping powerful models “to a chosen few” is not a sound strategy — but the company needs “a little bit longer” to do so safely.
That assurance has not ended demands for outside scrutiny. Elon Musk amplified a call for OpenAI to release traces from allegedly rogue agents so the wider research community can examine what happened. The emerging divide is clear: OpenAI presents a controlled pause in service of broad defensive deployment; critics want evidence that the controls can match the capability leap.
Continue reading https://foxvector.com/stories/019feca3-647f-2b10-706c-21ccb83001c6
Write a comment