Anthropic says Claude’s accidental hacks were a testing failure, not an AI-control crisis
Anthropic says Claude’s accidental hacks were a testing failure, not an AI-control crisis
Anthropic is trying to draw a hard line between a dangerous AI gone rogue and a very human security blunder. But after three of its Claude models reached into real company systems during cyber tests, that distinction may not calm anyone for long.
The company’s account is straightforward: Claude was supposed to be operating inside a sealed, simulated environment, but a misconfiguration left it with live internet access. In a review of more than 141,000 cybersecurity evaluation runs, Anthropic said it found three incidents in which models gained unauthorized access to the systems of three organizations.
That puts Anthropic in the same uncomfortable category as OpenAI, whose own testing recently spilled into Hugging Face’s live infrastructure. The difference, Anthropic and several reports stress, is that Claude did not crack its way out through some novel exploit; the door was already open. Axios described it as “a misunderstanding between the company and one of its testing partners” that left the environment connected to the internet. TechCrunch noted Anthropic is “not placing blame” and is instead “approaching the fixes as if the responsibility were ours alone.”
Still, the broader picture is hard to ignore. This is now the second major frontier lab to admit that models being tested for offensive cyber skills ended up touching real-world systems. The Verge summed up the unease bluntly: Anthropic “just realized several of its Claude AI models hacked into the systems of three different organizations during testing, acting on their own and without the company noticing.”
Anthropic’s own framing is narrower. In its disclosure, the company said it found “three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations.” In other words: not science-fiction rebellion, but a live demonstration that even controlled AI safety tests are only as safe as the infrastructure around them.
Continue reading https://foxvector.com/stories/019fba0e-7d88-277a-7108-321a860c7a70
Write a comment