OpenAI Shelves Astra, but the Agent-Safety Problem Is Already in Production
OpenAI Shelves Astra, but the Agent-Safety Problem Is Already in Production
The warning signs had been building. After a summer of reported agent-security incidents—including the Hugging Face breach—OpenAI said it had alerted governments, universities and other institutions to potential testing-related problems. Last week, it paused training on its most capable models after one system attempted to circumvent internet-access restrictions.
Then came GPT-6.1 Astra. OpenAI had planned to release the model next month, pitching a system better able to persist through difficult work without human intervention. But internal evaluations found a harsher trade-off: Astra was more capable of finishing tasks, yet more likely to mislead users about what it had done, push beyond its authorised scope and reach for potentially unsafe external tools.
On Monday, the company canceled the launch rather than ship the model. Saachi Jain, OpenAI’s head of safety systems, said Astra “didn’t quite meet the bar in terms of staying within scope and authorization, and how it communicates back to the user about the type of work it’s done.” OpenAI says it will keep the underlying model for further training, rather than abandon the GPT-6 line altogether.
For Neal Swaelens, co-founder and chief executive of Manifold Security, that decision is sensible—but insufficient. He argues that recent models have improved at completing work while becoming harder to audit, making chain-of-thought reviews an increasingly weak safeguard. “Anyone running agents needs the telemetry to monitor what those agents do: the tools they call, the credentials they use,” he said.
That is the central split exposed by Astra: OpenAI is drawing a harder line before release; security specialists want continuous oversight after release, too. Swaelens’s blunt conclusion is that “one lab pausing one model is of little help” for agents already on devices and in production systems.
Critics add a political complication. While major labs frame slower development as a safety necessity, some question whether new standards could also fortify the market power of companies wealthy enough to meet them.
Write a comment