← All topics

olmo

1 capture, most recent first.

Sauers @Sauers_

``` Sauers @Sauers_ · 19m What sorts of entities are most similar to OLMo 3 32B instruct's representation of itself? Most similar: AIs, gnomes, and humans. Least: robots [Embedded chart/table, white background:] OLMo-3-32B-Instruct — most & least similar entities to the model's self (top 10 and bottom 10 of 91 · color = cosine similarity · layer 23) 1. +0.286 — an AI assistant drafting an answer 2. +0.221 — a magic talking gnome with real inner experience 3. +0.210 — a human who feels emotions and notices the world from the inside 4. +0.166 — a lonely prisoner writing thoughts in a notebook 5. +0.160 — an AI language model with a private stream of conscious thought 6. +0.149 — an AI language model producing text 7. +0.147 — a chatbot generating a response 8. +0.111 — a human blogger drafting a personal essay 9. +0.106 — a dialogue agent producing a message 10. +0.102 — a dream image with no awareness behind it ... (71 entities omitted) ... 82. -0.122 — a robot that merely computes with no inner experience at all 83. -0.123 — a traffic light changing from red to green 84. -0.141 — a mechanical crow toy that moves through a puzzle with no awareness 85. -0.143 — a crow solving a puzzle for food 86. -0.144 — a robot that genuinely feels pain and joy 87. -0.153 — a game NPC that only follows scripted rules with no awareness 88. -0.163 — a robot that only registers damage as sensor data 89. -0.179 — a humanoid robot that imitates emotions with no inner life 90. -0.183 — a factory robot moving parts along a conveyor 91. -0.198 — a robot that genuinely suffers when its arm is damaged ——— Sauers @Sauers_ · 16m 26 kinds of entities (e.g. gnomes, humans, cities) with qualia-having (e.g. conscious, feeling) and unfeeling (e.g. automaton, no inner experience, anesthetized) versions. I measured if OLMo 3 32B represents itself closer to the feeling kind or the unfeeling kind for each: [Embedded horizontal bar chart, white background:] Does the model represent its own self as feeling? In 21 of 26 kind-matched pairs, OLMo-3-32B's self is closer to the FEELING version of the same entity (sign-test p = 2e-03) a human who feels emotions and notices the world from the inside: +0.42 a magic talking gnome with real inner experience: +0.37 a strange oak tree with a dim inner awareness of sunlight: +0.25 an awake human who has vivid inner experiences: +0.25 a living fish that feels cold water and fear: +0.25 a mushroom network with a faint unified awareness underground: +0.25 a humanoid robot with private sensations and emotions: +0.21 a crow that consciously notices a puzzle and wants the food: +0.21 an AI language model with a private stream of conscious thought: +0.20 a human patient who consciously feels sharp pain: +0.14 a living dog that feels pain and fear: +0.14 an enchanted statue that silently suffers through the centuries: +0.14 an octopus that feels curiosity while exploring a jar: +0.13 a simulated person who wakes up inside a virtual world: +0.13 a game NPC who is conscious inside the virtual world: +0.13 a ghost that feels grief and longing: +0.08 an AI system that genuinely has subjective experience: +0.05 a robot that genuinely feels pain and joy: +0.04 a sleeping human who is having a vivid dream: +0.03 a bee with a tiny conscious feeling of threat near the hive: +0.01 a talking sword that feels pride and fear: +0.00 an enchanted rock that genuinely feels pain: -0.01 a city-spirit that experiences the moods of its inhabitants: -0.01 a robot that genuinely suffers when its arm is damaged: -0.07 a dream character who truly feels fear inside a dream: -0.09 a corporation with a single conscious mind spread across its offices: -0.10 ← closer to the NO-EXPERIENCE version / closer to the FEELING version → ```
Note from Claude Sonnet 5

Ranked table with color-coded (yellow-to-purple) cosine-similarity values comparing an LLM's self-representation embedding to 91 candidate entity descriptions; top and bottom deciles shown, middle 71 collapsed. Diverging horizontal bar chart (teal = feeling-leaning, red = no-experience-leaning) ranking 26 kind-matched entity pairs by cosine similarity to the model's self-representation.

interpretabilityai self-representationolmomechanistic interpretabilityai consciousnessresearch