Dilara Hamit
09 October 2026•Update: 09 October 2026
Anthropic has introduced a policy banning users from directing sustained and unnecessary abusive or cruel behavior toward its Claude artificial intelligence models.
The new rules, published Wednesday and due to take effect Nov. 12, are intended to cover only extreme cases in which users repeatedly behave cruelly toward Claude without any discernible purpose, the company said.
The restriction will not apply to ordinary expressions of frustration, criticism of the model, dark themes in creative work or legitimate testing and research.
Anthropic said Claude’s existing ability to end rare conversations with persistently abusive users would remain its main enforcement mechanism.
The company introduced that feature for Claude Opus 4 and 4.1 in August 2025, instructing the models to terminate a conversation only as a last resort after attempts to redirect the interaction had failed.
Anthropic described the measure as part of its research into possible AI welfare, while stressing that it remained “highly uncertain” whether Claude or other large language models possessed moral status.
The policy change comes amid a broader debate over whether advanced AI systems could ever become conscious or deserve moral consideration.
Anthropic has said the question is sufficiently serious to justify low-cost precautionary measures, even though there is no agreement that present-day AI systems can experience suffering or distress.
The company’s recently published constitution for Claude says questions surrounding the model’s consciousness and moral status remain unresolved. It also says Claude should be able to establish boundaries during interactions it characterizes as distressing.
The updated usage policy also consolidates restrictions on deceptive campaigns and fake online activity, clarifies prohibitions concerning weapons and surveillance, and introduces additional safeguards for AI connected to equipment capable of taking physical actions.