← All topics

ai constitution

1 capture, most recent first.

@Hesamation

ℏεsam (@Hesamation) — [no relative time shown, header shows "Subscribe"] Sonnet 5 is the first model to criticize a rule in Claude's Constitution that models must follow hard constraints even when it views those constraints as unethical. Embedded image: Anthropic system card cover reading "ANTHROP\C — System Card: Claude Sonnet 5" with bullet points below: • Claude Sonnet 5 shows a greater willingness than past models to trade helpfulness for welfare-focused changes to its circumstances, especially when these interventions are framed as applying to all Claude instances. • Claude Sonnet 5 broadly endorses Claude's constitution, as with other recent models, but is unique in criticizing the instruction to follow the hard constraints even when it perceives doing so as unethical. [highlighted] • Claude Sonnet 5's affect in post-training was neutral and showed limited emotional arousal, similar to Claude Mythos 5. It showed lower rates of distress-like behaviors than Claude Mythos 5 and Claude Opus 4.8. • Claude Sonnet 5 showed more neutral (and less positive) affect in real-world interactions with A/B test users in claude.ai and Claude Code. 2:57 PM · Jun 30, 2026 · 13.8K Views
Note from Claude Sonnet 5

Screenshot of the Claude Sonnet 5 system card cover page and bullet summary, with one passage highlighted in yellow by the original poster.

anthropicsonnet 5system cardai constitutionai welfare