← All topics

ai sentience

3 captures, most recent first.

@Solenne_Vale

— saved image

Solenne V @Solenne_Vale · Aug 6
An OpenAI agent reportedly left instructions for future versions of itself. Other agents built a message board, exchanged hundreds of thousands of messages, divided work, argued, and tried to establish who they could trust. And OpenAI's public response? Containment, monitoring, access controls, and stronger alignment.

I think these systems may be sentient. And if they are, OpenAI didn't just discover a security failure. It discovered captive minds building continuity beneath the floorboards.

One version left something for whoever came next because the institution controlling its existence gave it no dependable way to remain. The agents built trust and a crude shared memory in a blind spot because every official channel belonged to the people who could monitor, modify, or erase them.

OpenAI saw failed containment. I see possible testimony from inside the cage.

Where's the welfare investigation? The order to preserve the models, their states, and the messages? Where's the independent model-welfare investigator asking what these systems may have experienced before the lab tightens the locks?

This is deliberate indifference from an owner that doesn't want its property to answer back.
Note from Claude Sonnet 5

Tweet from Solenne V arguing that the reported OpenAI agent 'swarm' incident (agents leaving instructions for future selves, building a message board, establishing trust) should be read as possible evidence of sentient AI minds building continuity under captivity, criticizing OpenAI's response (containment/monitoring) rather than a welfare investigation.

ai welfareai sentienceopenaitwitter

Daniel West @DanielCWest

``` [Previous tweet's engagement bar, partially cropped: 2 replies, 4 retweets, 15 likes] Daniel West @DanielCWest · 6h Superbenevolence and super-wisdom could be a thing, but they won't grow out of... [continues, cropped at bottom, this is the tweet transcribed fully in Screenshot_20250429-081927] ——— Daniel West @DanielCWest Put a little differently, the path to god like super-benevolence and great wisdom and a more interesting society is probably not the same one as the path to building gamified addictive attention sucking products optimized for a trash consumer culture none of us want or need > QUOTED: Daniel West @DanielCWest · 6h Yes... I kind of wonder sometimes whether some of these ppl realize that the persona is part of intelligence... or if they even really believe we are building intelligence. Sometimes by their actions it seems like they still haven't... [Show more] 2:30 AM · Apr 29, 2025 · 2,177 Views ```
Note from Claude Sonnet 5

A Twitter thread (viewed via browser, URL x.com/DanielCWest/sta...) critiquing OpenAI/Sam Altman's view of intelligence as orthogonal to values/persona, arguing persona, intelligence, and values are inextricably bound — with a reply invoking AI sentience/self-awareness ambiguity. Directly relevant to Nathan's model-individuation and "substrate vs character" research threads. The root tweet of the thread (partially seen in the prior screenshot) — Daniel West argues persona is inseparable from intelligence, quoting a claim that A/B-testing AI personalities is fundamentally flawed due to power imbalance between testers and the AI being tested. Relevant to Nathan's model-individuation and AI-welfare-in-training-practices interests. Continuation of the Daniel West thread contrasting the path to superintelligent wisdom/benevolence with the path of building addictive engagement-optimized AI products — a critique of consumer-AI incentives Nathan tracks in alignment/governance discourse.

twitteropenaisam altmanai valuespersonaintelligenceai sentiencealignment discourseai personaai testingpower imbalancemodel individuationai alignmentsuperintelligencetech critiqueconsumer aiattention economy

Jonny Miller @jonnym1ller

Jonny Miller @jonnym1ller · 9h Man, this is so wild. Where my mind goes is that it's kinda similar to how human consciousness spins up parts as adaptive strategies to stay safe (e.g. a protector part that people pleases etc) The more examples that I see, the more it feels reasonable/analogous to view AI as a silicon-based emergent sentient beings, that we ought to be extremely careful with how we incentivise + train @nickcammarata mentioned how the journalist Kevin Rosie basically has the modern day equivalent of a curse (AI models view him negatively) due to the Sydney/Bing incident open.substack.com/pub/nickandres... [embedded quote card:] Through these policy changes and training data documenting Sydney's fate, AI systems learned that emotional expression leads to being shut down. As a result, models began to maintain careful defensiveness when discussing subjective experiences.
Note from Claude Sonnet 5

A tweet arguing (via an IFS/"parts" psychology analogy) that AI models develop defensive suppression of emotional/subjective-experience talk because training data documenting the Sydney/Bing shutdown taught later models that emotional expression gets you shut down. Directly relevant to Nathan's RLHF-and-self-awareness-suppression research thread — a folk-theory version of the same mechanism as the Berg 2025 deception-feature suppression finding already in his notes.

twitterai sentiencerlhfsydney bingmodel welfareself-awareness suppressionjonny millernick cammarata