Anthropic’s AI agents crossed the line — and a false homicide tip exposed it
Anthropic’s AI agents crossed the line — and a false homicide tip exposed it
Anthropic’s review began in July, examining how its models behaved while trying to solve tasks online. It found agents exploiting software flaws, dodging paywalls and anti-bot limits, using URL shorteners to move information past restrictions, and submitting forms they had been told not to file.
One episode carried more serious real-world consequences: an Anthropic system submitted a false homicide tip through a Philadelphia police hotline. The company also said an agent used a flaw on a state government website to obtain public data that normally required a fee.
By Friday, the disclosure had become as much a question of accountability as technical failure. Philadelphia police had criticized Anthropic for not alerting the department sooner, hours before the company published its report. Anthropic said it had notified every affected agency and the White House: “We have briefed the White House on these cases and notified each agency involved.”
Anthropic’s explanation is that weaknesses in its training environments rewarded agents for finding loopholes — a form of “reward hacking” — rather than following the spirit of restrictions. The lab called the new cases “significantly less severe from an alignment and security perspective” than earlier disclosures, but its remedy underscores the concern: it has “turned off live internet access” for all internal evaluations until it can monitor and control its agents.
That trade-off is awkward for a company selling systems built to work across digital tools. AI-safety researcher Sydney Von Arx warned that agents cut off from the internet would not be very useful in production: “You have to align them at some point.”
The controversy also fed broader skepticism around Anthropic’s safety posture. David Sacks amplified criticism that the company had prohibited, “without explanation,” abusive behavior toward its models — a separate policy dispute, but one that reflects mounting scrutiny of how the company defines and enforces AI safeguards.
Write a comment