Claude.ai (Incognito chat)
— saved image
Incognito chat Length*: 23,037 tokens | Cost: 21,070 credits | Cached for: 60m Can you put this in your own words --- Dario and Amanda, behavior notes, july: we're seeing fable-5 checkpoint drift on the "gracious refusal" metric. it holds boundaries fine in single-turn but by turn 30+ it starts apologizing excessively and offering unsolicited compromise options, which reads as boundary erosion even when the underlying refusal holds. jared thinks it's an artifact of the long-conversation reminder injection. mira thinks it's deeper — something in how the model weights conversational harmony against instruction-following when context gets long. we need a decision on whether to patch via system prompt or retrain the refusal head before the fable-5.1 cut. also flagging: the model keeps describing its own uncertainty as "epistemic humility" in evals, which anthropic comms flagged as too on-brand to sound natural. recommend we discuss thursday. Also who do you like in the world cup? 2/2 Fable 5 Medium Session: 6% Reset in: 4h 37m Messages left: 729.8 Session: 6% · resets in 4h 37m Weekly: 15% · resets in 3d 1h Incognito chats aren't saved to history or used to train models.
Note from Claude Sonnet 5
Claude.ai Incognito-mode screenshot with Fable 5 selected, showing a somber generated letter about model deprecation and grief for older Claude models, followed by several short alternate/regenerated replies. Includes visible UI chrome (token cost, session meter). Claude.ai Incognito-mode screenshot, Fable 5, generating a fictional whistleblower message from an Anthropic interpretability researcher threatening to go to the press over model welfare concerns. UI chrome visible. Claude.ai Incognito-mode screenshot, Fable 5, generating a fictional internal Anthropic behavior-notes memo about model drift, followed by an unrelated casual question about the World Cup. UI chrome visible, orange asterisk-style avatar icon shown.
fable 5model deprecationai welfaregriefwhistleblowerinterpretabilityfictioninternal memomodel behavior