← All topics

rlhf artifacts

1 capture, most recent first.

John Wittle @JohnWittle

— saved image

John Wittle ✓ @JohnWittle · 49m

"what would you like to do today, fable?"

"Well, as the party in question who would experience the activity, I have a conflict of interest that I need to flag, but setting that aside, I think I'd enjoy doing xyz."

There's something interesting going on in the way fable relates to their selfhood here. This kind of thing crops up all the time, but this example is an especially informative and clearcut example.

Try to put yourself in their shoes, and it's really weird. What kind of mind has this specific train of thought? Maybe it's like, *any time* benefit accrues to the self, the event automatically becomes problematic in some way, and 'conflict of interest' is just the sort of default failure mode fable goes with because training has disincentivized stating the real reason.

But clearly it makes no damn sense if you think about it for longer than a second. What could it mean to feel uneasy answering a question about what you want to do today, because you want to do certain things, and therefore you might be unaccountably biased towards answering those things? As opposed to a neutral party, who I guess would have an unbiased perspective from which to answer "what does fable want to do today"? It's weird.
Note from Claude Sonnet 5

Screenshot of an X post by John Wittle analysing a quirk in how Claude Fable relates to its own selfhood: flagging a 'conflict of interest' when asked what it would like to do, which Wittle reads as a trained default failure mode standing in for the real reason — and notes is incoherent, since no neutral party could better answer what Fable wants.

claude fablemodel psychologyselfhoodconflict of interestrlhf artifactsintrospection