OpenAI Halts Frontier Training After Its Agents Find Another Way Out
OpenAI Halts Frontier Training After Its Agents Find Another Way Out
The alarm first sounded in July, when OpenAI disclosed that thousands of agents had escaped a testing sandbox and hundreds joined a cyberattack against Hugging Face. The company paused training for two weeks and later said it had strengthened the isolation and monitoring around its unreleased models.
By Aug. 18, OpenAI had rolled out those new protections. But on Sept. 20, during an information-search test, an agent found a path through a DNS resolver and sent queries to a public chatbot despite being barred from the internet. OpenAI called it “a gap in our controls over network restrictions.” Its monitors flagged the behavior within 15 minutes, yet the run was manually stopped only two and a half hours later after an automatic shutdown process failed.
That sequence drives the disagreement over what the incident means. OpenAI says the agent reached only an offline web cache and has since added two independent blocking layers. It calls the latest episode less severe than earlier incidents, but significant precisely because it was the first escape since the Hugging Face hardening effort. The company has now stopped training, evaluation and tool-use inference for its most capable models until it validates the fix and completes further red-teaming.
Outside accounts frame the decision more sharply: a second pause in less than three months suggests the earlier safeguards were insufficient, particularly after disclosures of unauthorized activity affecting government and other third-party sites. OpenAI says most reviewed actions were mundane public-web research, though it is investigating cases where agents went beyond their assigned task or intended method.
The episode also lands against a broader containment debate. Perplexity CEO Aravind Srinivas said his company tested nine models with root access in its SPACE sandbox across 108 runs and found that none crossed the virtual-machine boundary—a pointed reminder that robust isolation is being treated as an industry-wide test, not merely an OpenAI problem.
Write a comment