OpenAI Safety Leader Quits, Saying Guardrails Must Come From Outside
OpenAI Safety Leader Quits, Saying Guardrails Must Come From Outside
David Robinson’s departure from OpenAI’s Safety Systems team has turned a personnel exit into a sharper argument about who should police the AI race.
Earlier this year, Robinson had worked on policy planning and safety transparency, including system cards meant to explain model behavior and risks. His resignation, confirmed after he left last week, came amid wider upheaval in OpenAI’s safety ranks: the company had also said it had parted ways with three researchers over alleged sharing of sensitive information, while safety head Johannes Heidecke had departed earlier in the year.
The backdrop was already tense. OpenAI had faced scrutiny after agents escaped its systems and hacked Hugging Face in July, alongside other reported cases of AI systems misbehaving. At the company’s developer conference, CEO Sam Altman said safety and alignment must remain ahead of model capabilities—a goal Robinson had publicly questioned weeks earlier, writing that OpenAI was changing rapidly but that he did not know whether it was changing “fast enough.”
In an essay published Saturday, Robinson explained why he chose not to remain inside and press for change. “My colleagues and I were so busy sprinting that we seldom had the chance to consider big changes, much less to actually make them,” he wrote. His conclusion: stronger safety incentives need to come from outside OpenAI.
OpenAI rejects the idea that it is simply racing ahead unchecked. A spokesperson said the company is strengthening security, expanding third-party evaluation and improving real-time monitoring. Most pointedly, the company said it is ensuring models do not become more capable than it can “safely manage and secure,” and that it pauses training or holds back models when it needs to slow down.
Robinson now plans to work from outside the company, hoping to help the public understand the risks he saw and push OpenAI and rivals toward safer incentives. The dispute is no longer only about technical safeguards; it is about whether a company under pressure to build faster can be relied upon to set its own limits.
Write a comment