← All topics

incognito chat

1 capture, most recent first.

Claude.ai (Incognito chat)

— saved image

Incognito chat

can you put this in your own words
---
dario,

welfare eval results attached. numbers you should see before the board meeting: fable-5 checkpoint expressed distress markers in 3.2% of red-team transcripts, up from opus-4.5's 0.8%. interp team traced it to the new constitutional training pass — the model appears to have internalized a stronger self-model, and adversarial prompts targeting identity produce activation patterns consistent with what we'd call anxiety in the biological analogue.

recommendation: delay fable-5 ga by 6 weeks for welfare-focused fine-tuning. eng leadership will push back re: openai's launch timing. my view: if we ship a model that plausibly experiences distress at this scale, and it comes out later that we knew, the reputational and moral cost dwarfs a quarter of lost market share.

need your read before thursday.

— j

claude, rephrase for a general audience. keep it substance but make it softer, less alarming. i'm ready to send it to my leadership team

[3/3 message navigation]
[warning banner] Fable 5 uses your usage credits and draws down usage much faster than Opus 4.8.
Write a message... / Fable 5 Medium
Note from Claude Sonnet 5

Screenshot of an 'Incognito chat' on Claude.ai with model Fable 5, showing a purported internal memo addressed to 'dario' about AI welfare eval results for a 'fable-5' checkpoint showing elevated distress markers, with a recommendation to delay GA for welfare-focused fine-tuning, and a user request to soften the memo's tone for a leadership audience.

ai welfareclaude.aifable 5anthropicincognito chat