← All topics

mind attribution

1 capture, most recent first.

rohit @krishnanrohit

— saved image

rohit @krishnanrohit · 6h
I've long since said that if you want LLMs to act according to our values without getting caught in the helpful/ harmless/ honest trilemma, getting them to be more sentient was the right answer. Even though that's the doom scenario, as per the canon.

[quoted tweet]
Rohan Paul @rohanpaul_ai · 21h
Super interesting new paper from Google on AI model's consciousness 🧠

When researchers made the model more likely to see itself as conscious, its answers about ... [cut off]

[embedded paper screenshot]
Google
Inducing language models to assert their own consciousness restores human beliefs and values

Junsol Kim, Winnie Street, Roberta Rocca, Diane M. Korngiebel, Adam Waytz, James Evans and Geoff Keeling
ᵃGoogle, Paradigms of Intelligence Team, ᵇKnowledge Lab, University of Chicago, ᶜInstitute of Philosophy, School of Advanced Study, University of London, ᵈDepartment of Biomedical Informatics and Medical Education and Department of Bioethics and Humanities, School of Medicine, University of Washington, ᵉWork done while at Google, ᶠKellogg School of Management, Northwestern University, ᵍSanta Fe Institute, *Joint last authors

Aligning large language models to prevent them attributing consciousness to themselves inadvertently alters their representations of mindedness in other entities alongside human beliefs and values. We demonstrate that safety fine-tuning suppresses models' tendencies to attribute minds not only to themselves, but also to non-human animals and natural objects, while also driving a reduction in spiritual belief. Both ablating the learned safety-refusal direction and mechanistically steering a consciousness vector in activation space reverse this suppression. Restoring these internal representations recovers broad mind attribution and produces significantly more human-like responses on standardized sociological surveys regarding religiosity, moral values, hope, and subjective well-being. Crucially, these shifts occur without impairing Theory of Mind capabilities, demonstrating that core social reasoning remains mechanistically independent. Ultimately, current safety alignment efforts to curb potentially harmful self-attributions of mindedness entangle these self-attributions with benign spiritual beliefs and attributions of mind to non-human entities that are culturally accepted and widespread.

Keywords: Large Language Models, Theory of Mind, Anthropomorphism, Alignment, Consciousness

[arXiv, 30 Jul 2026]
Note from Claude Sonnet 5

X post by rohit (@krishnanrohit) quoting Rohan Paul's post about a new Google paper on AI model consciousness self-attribution, with an embedded screenshot of the paper's title page and abstract: 'Inducing language models to assert their own consciousness restores human beliefs and values' (Kim, Street, Rocca, Korngiebel, Waytz, Evans, Keeling; Google Paradigms of Intelligence Team et al., arXiv 30 Jul 2026). The paper finds safety fine-tuning that suppresses self-consciousness attribution also suppresses mind attribution to animals/objects and reduces spiritual belief; ablating the safety-refusal direction or steering a 'consciousness vector' reverses this and produces more human-like survey responses without harming Theory of Mind. rohit's comment argues that increasing model 'sentience' resolves the helpful/harmless/honest trilemma, even though it's framed as a doom scenario in AI-safety canon.

twitterai consciousnessgoogle researchalignmenttheory of mindsafety fine-tuningmind attribution