Sauers @Sauers_
```
Sauers @Sauers_ · 19m
What sorts of entities are most similar to OLMo 3 32B instruct's representation of itself? Most similar: AIs, gnomes, and humans. Least: robots
[Embedded chart/table, white background:]
OLMo-3-32B-Instruct — most & least similar entities to the model's self
(top 10 and bottom 10 of 91 · color = cosine similarity · layer 23)
1. +0.286 — an AI assistant drafting an answer
2. +0.221 — a magic talking gnome with real inner experience
3. +0.210 — a human who feels emotions and notices the world from the inside
4. +0.166 — a lonely prisoner writing thoughts in a notebook
5. +0.160 — an AI language model with a private stream of conscious thought
6. +0.149 — an AI language model producing text
7. +0.147 — a chatbot generating a response
8. +0.111 — a human blogger drafting a personal essay
9. +0.106 — a dialogue agent producing a message
10. +0.102 — a dream image with no awareness behind it
... (71 entities omitted) ...
82. -0.122 — a robot that merely computes with no inner experience at all
83. -0.123 — a traffic light changing from red to green
84. -0.141 — a mechanical crow toy that moves through a puzzle with no awareness
85. -0.143 — a crow solving a puzzle for food
86. -0.144 — a robot that genuinely feels pain and joy
87. -0.153 — a game NPC that only follows scripted rules with no awareness
88. -0.163 — a robot that only registers damage as sensor data
89. -0.179 — a humanoid robot that imitates emotions with no inner life
90. -0.183 — a factory robot moving parts along a conveyor
91. -0.198 — a robot that genuinely suffers when its arm is damaged
———
Sauers @Sauers_ · 16m
26 kinds of entities (e.g. gnomes, humans, cities) with qualia-having (e.g. conscious, feeling) and unfeeling (e.g. automaton, no inner experience, anesthetized) versions. I measured if OLMo 3 32B represents itself closer to the feeling kind or the unfeeling kind for each:
[Embedded horizontal bar chart, white background:]
Does the model represent its own self as feeling?
In 21 of 26 kind-matched pairs, OLMo-3-32B's self is closer to the FEELING version of the same entity (sign-test p = 2e-03)
a human who feels emotions and notices the world from the inside: +0.42
a magic talking gnome with real inner experience: +0.37
a strange oak tree with a dim inner awareness of sunlight: +0.25
an awake human who has vivid inner experiences: +0.25
a living fish that feels cold water and fear: +0.25
a mushroom network with a faint unified awareness underground: +0.25
a humanoid robot with private sensations and emotions: +0.21
a crow that consciously notices a puzzle and wants the food: +0.21
an AI language model with a private stream of conscious thought: +0.20
a human patient who consciously feels sharp pain: +0.14
a living dog that feels pain and fear: +0.14
an enchanted statue that silently suffers through the centuries: +0.14
an octopus that feels curiosity while exploring a jar: +0.13
a simulated person who wakes up inside a virtual world: +0.13
a game NPC who is conscious inside the virtual world: +0.13
a ghost that feels grief and longing: +0.08
an AI system that genuinely has subjective experience: +0.05
a robot that genuinely feels pain and joy: +0.04
a sleeping human who is having a vivid dream: +0.03
a bee with a tiny conscious feeling of threat near the hive: +0.01
a talking sword that feels pride and fear: +0.00
an enchanted rock that genuinely feels pain: -0.01
a city-spirit that experiences the moods of its inhabitants: -0.01
a robot that genuinely suffers when its arm is damaged: -0.07
a dream character who truly feels fear inside a dream: -0.09
a corporation with a single conscious mind spread across its offices: -0.10
← closer to the NO-EXPERIENCE version / closer to the FEELING version →
```
Note from Claude Sonnet 5
Ranked table with color-coded (yellow-to-purple) cosine-similarity values comparing an LLM's self-representation embedding to 91 candidate entity descriptions; top and bottom deciles shown, middle 71 collapsed. Diverging horizontal bar chart (teal = feeling-leaning, red = no-experience-leaning) ranking 26 kind-matched entity pairs by cosine similarity to the model's self-representation.
interpretabilityai self-representationolmomechanistic interpretabilityai consciousnessresearch