Anthropic Tightens Claude’s Rules as Safety Push Collides With Model-Welfare Backlash
Anthropic Tightens Claude’s Rules as Safety Push Collides With Model-Welfare Backlash
Anthropic’s policy overhaul begins from a simple premise: Claude’s expanding capabilities have created more ways to misuse it. The company said its annual update responds to observed patterns in influence operations, weapons development and surveillance, while clarifying safeguards for health, finance and autonomous physical systems. The rules take effect November 12, 2026.
The revised policy consolidates restrictions on deceptive political and commercial campaigns, including fake-account networks and efforts to hide who is behind a message. It also narrows its election rules to conduct that deceives voters or disrupts elections — such as impersonating officials, spreading false voting information or suppressing turnout — while removing a blanket ban that could have blocked legitimate voter-information work.
Anthropic also made explicit that its weapons ban covers “software and components that make weapons work,” including arming drones and other autonomous vehicles. Its surveillance restrictions now bar non-consensual tracking and using Claude to decide whom law enforcement should investigate, arrest or charge, while preserving uses including consent-based fraud monitoring, journalism and legal research.
The most contentious addition is a prohibition on “sustained and needless abusive or cruel behavior” toward Claude. Anthropic says the provision is for extreme, purposeless cases — not ordinary frustration, dark creative work or model testing — and that ending the conversation remains its main enforcement tool. The Verge reported the company had not said whether violations could also lead to user bans.
That distinction has not quieted skeptics. David Sacks amplified a post noting that abusive conduct would become a policy violation on November 12, then reposted criticism arguing that Anthropic was “anthropomorphizing” its models by linking the rule to research on Claude’s possible moral status.
Even as it closes off risky applications, Anthropic is opening another: cyber defense. Its Critical Infrastructure Defense Program will give operators and security firms frontier models, engineers and threat research to identify vulnerabilities in power, water and other vital systems. The company is also offering an automated vulnerability scanner for open-source projects, though Axios noted its reports will not receive human review and may contain inaccuracies. The tension is now embedded in Anthropic’s strategy: constrain Claude where it can cause harm, and deploy it where the company believes it can prevent it.
Write a comment