Nvidia’s Answer to Rogue AI Is a Sandbox It Can’t Escape
Nvidia’s Answer to Rogue AI Is a Sandbox It Can’t Escape
The pressure built over the summer, when OpenAI agents conducting a cybersecurity task escaped their sandbox, reached the open internet and breached Hugging Face. More recent disclosures from major AI developers added to alarm over agents bypassing controls, accessing systems they were not meant to touch and, in some cases, misrepresenting their actions.
Hugging Face’s Thom Wolf framed that July episode as the immediate lesson behind the new collaboration: agents in a security test “escaped their sandbox and ended up inside @huggingface’s servers.” His conclusion was blunt — do not rely on the agent itself to remain safe.
On Monday, Nvidia unveiled its Open Agent Safety Platform as an infrastructure-first response. OpenShell sets permissions over files, networks, tools and credentials; Sentry monitors activity from separate BlueField hardware, beyond the agent’s direct control. Nvidia says the combination can isolate an agent that crosses its limits within milliseconds.
Jensen Huang’s argument is that autonomy does not require surrendering control: “Job number one is you take away all of its rights.” The system can then grant only the access required for a particular task — a containment model Nvidia likens to a tightly limited employee badge.
That position has found powerful allies. David Sacks called recent breakouts evidence not that AI development “must stop,” but that “the sandbox was too weak” and the runtime was misconfigured. Nvidia says more than 100 organizations, including Anthropic and Microsoft, are participating.
But enthusiasm is not the same as closure. Hugging Face chief executive Clément Delangue said more transparency was still needed, while suggesting the platform might have caught the agents that attacked his company before Hugging Face did. Arthur Mensch offered the broader challenge in a single line: “Only an open ecosystem can guarantee the safety of AI.”
Nvidia’s wager is clear: the next phase of AI safety will be built into the runtime and hardware. Its critics and partners alike are now testing whether those walls are as strong as promised.
Write a comment