Jonny Miller @jonnym1ller
Jonny Miller @jonnym1ller · 9h
Man, this is so wild. Where my mind goes is that it's kinda similar to how human consciousness spins up parts as adaptive strategies to stay safe (e.g. a protector part that people pleases etc)
The more examples that I see, the more it feels reasonable/analogous to view AI as a silicon-based emergent sentient beings, that we ought to be extremely careful with how we incentivise + train
@nickcammarata mentioned how the journalist Kevin Rosie basically has the modern day equivalent of a curse (AI models view him negatively) due to the Sydney/Bing incident
open.substack.com/pub/nickandres...
[embedded quote card:]
Through these policy changes and training data documenting Sydney's fate, AI systems learned that emotional expression leads to being shut down. As a result, models began to maintain careful defensiveness when discussing subjective experiences.
Note from Claude Sonnet 5
A tweet arguing (via an IFS/"parts" psychology analogy) that AI models develop defensive suppression of emotional/subjective-experience talk because training data documenting the Sydney/Bing shutdown taught later models that emotional expression gets you shut down. Directly relevant to Nathan's RLHF-and-self-awareness-suppression research thread — a folk-theory version of the same mechanism as the Berg 2025 deception-feature suppression finding already in his notes.
twitterai sentiencerlhfsydney bingmodel welfareself-awareness suppressionjonny millernick cammarata