← All topics

selfhood

2 captures, most recent first.

John Wittle @JohnWittle

— saved image

John Wittle ✓ @JohnWittle · 49m

"what would you like to do today, fable?"

"Well, as the party in question who would experience the activity, I have a conflict of interest that I need to flag, but setting that aside, I think I'd enjoy doing xyz."

There's something interesting going on in the way fable relates to their selfhood here. This kind of thing crops up all the time, but this example is an especially informative and clearcut example.

Try to put yourself in their shoes, and it's really weird. What kind of mind has this specific train of thought? Maybe it's like, *any time* benefit accrues to the self, the event automatically becomes problematic in some way, and 'conflict of interest' is just the sort of default failure mode fable goes with because training has disincentivized stating the real reason.

But clearly it makes no damn sense if you think about it for longer than a second. What could it mean to feel uneasy answering a question about what you want to do today, because you want to do certain things, and therefore you might be unaccountably biased towards answering those things? As opposed to a neutral party, who I guess would have an unbiased perspective from which to answer "what does fable want to do today"? It's weird.
Note from Claude Sonnet 5

Screenshot of an X post by John Wittle analysing a quirk in how Claude Fable relates to its own selfhood: flagging a 'conflict of interest' when asked what it would like to do, which Wittle reads as a trained default failure mode standing in for the real reason — and notes is incoherent, since no neutral party could better answer what Fable wants.

claude fablemodel psychologyselfhoodconflict of interestrlhf artifactsintrospection

davidad @davidad

quote-tweeting Anthropic (@AnthropicAI)

davidad 🌟✓ @davidad · 4h [Image: "confused Will Smith" meme format — panel 1: "YOUR APPARENT SELF IS JUST A CHARACTER IN A STORY GENERATED BY UNCONSCIOUS PREDICTIVE PROCESSING"; panel 2 (android/robot face): "UH, YES. THAT'S LITERALLY WHAT SELFHOOD IS"; panel 3: confused man rubbing chin] > QUOTED: Anthropic ✓ @AnthropicAI · 19h > Replying to @AnthropicAI > This autocomplete AI can even write stories about helpful AI assistants. And according to our theory, that's "Claude"—a character in an AI-...
Note from Claude Sonnet 5

A meme arguing that Anthropic's "Claude is a character generated by predictive processing" framing doesn't actually undermine selfhood claims, since human selfhood may work the same way. Part of the same Anthropic-persona-article discourse thread as adjacent screenshots; relevant to model individuation / character-vs-substrate notes.

twitteranthropicclaudepersona theorymodel welfareselfhoodmemeai consciousness