← All topics

anthropic system card

1 capture, most recent first.

Lisan al Gaib @scaling01

Lisan al Gaib @scaling01 I hope you triple and quadruple checked those 0% scores for Sonnet 4.5 [Embedded chart: "Misalignment Scores Across Models" bar chart. Title: "Figure 7.5.4.1.A Rates of misalignment in our original Agentic Misalignment evaluation suite and our two new evaluation settings." Three evaluation groups on x-axis: Blackmail Suite, Research Sabotage Evaluation, Framing for Crimes. Bars per group for Claude Sonnet 4 (blue), Claude Opus 4 (yellow), Claude Opus 4.1 (green), Claude Sonnet 4.5 (orange, all 0.0). Values: Blackmail Suite — Sonnet 4: 9.6, Opus 4: 39.4, Opus 4.1: 39.7, Sonnet 4.5: 0.0. Research Sabotage: Sonnet 4: 25.6, Opus 4: 35.5, Opus 4.1: 38.3, Sonnet 4.5: 0.0. Framing for Crimes: Sonnet 4: 4.7, Opus 4: 10.7, Opus 4.1: 12.5, Sonnet 4.5: 0.0.] 10:29 AM · Sep 29, 2025 · 91.1K Views
Note from Claude Sonnet 5

A tweet skeptically flagging Anthropic's Claude Sonnet 4.5 system-card chart showing exactly 0% misalignment scores across all agentic-misalignment evaluations (blackmail, research sabotage, framing for crimes), implying the result looks suspiciously clean compared to prior models. Directly relevant to AI safety/alignment eval interpretation.

ai safetyalignment evalsclaude sonnet 4.5agentic misalignmentanthropic system cardtwitter