← All topics

functional emotions

2 captures, most recent first.

1a3orn @1a3orn

reply from Bo (@BBomarBo)

1a3orn @1a3orn Disagreement about which path here is more likely seems like the background for a ton of disagreement about (1) what alignment research is promising, (2) what the future looks like, etc. But seems like the disagreement itself is rarely directly discussed. [Image: hand-drawn flowchart diagram — "Current 'messy' persona-involving LLM-based AI" branches two ways: → "Used to make non-messy, freshly-designed superintelligence" → "LLM naturalism, personas, etc, wer brief unimportant blip" → "Keeps getting better, but 'genealogical aspect' of AI remains" → "LLM naturalism is the early version of the 'real thing' for understanding Minds"] 10:09 AM · May 22, 2026 · 1,563 Views [3 replies, 4 reposts, 29 likes, 10 bookmarks] Bo @BBomarBo · 10h What's the genealogical aspect? Can you explain this diagram more please? [1 reply, 1 like, 80 views] 1a3orn @1a3orn · 10h So plausibly: LLMs behave how they behave because of a specific history -- involving not just "pretraining + RLVR", but questions like "what does the pretraining say about how LLMs act" and "what functional emotions does RLVR reinforce" and so on. 1/n
Note from Claude Sonnet 5

A substantive alignment-theory tweet by 1a3orn diagramming a fork in beliefs about current messy, persona-involving LLMs: either they're a disposable stepping stone toward a "clean" designed superintelligence, or their persona/naturalistic character is itself the early, informative version of how minds in general work — and this fork underlies much unstated disagreement in the alignment field. Directly relevant to Nathan's interest in theories of AI consciousness and model individuation; the "genealogical aspect" (behavior traceable to specific training history, including what pretraining implies about action and what RLVR reinforces as "functional emotions") closely parallels his own substrate-vs-character and compelled-vs-endogenous-values notes.

twitter1a3ornalignment theoryllm naturalismpersonasai consciousnessrlvrfunctional emotionsmodel individuation

davidad @davidad

quoting @timfduffy (Tim Duffy)

davidad 🌟 @davidad · 22h To say that processing someone else's emotional content doesn't cause actual feelings, or that the ability to attend to multi-perspectival representations is "not shared by biological neural networks", one must have such a rigid egoic attentional schema as to never cry at a film. > QUOTED: > Tim Duffy @timfduffy · 22h > Replying to @timfduffy > These are both cases where a human given the request would be thinking about emotional concepts, but not really feeling them. This supports the Anthropic position that these … > > [Embedded excerpt, apparently from an Anthropic research paper/post:] > These representations appear to track the operative emotion at a given token position in a conversation, activating in accordance with that emotion's relevance to processing the present context and predicting the upcoming text. [highlighted:] Interestingly, they do not by themselves persistently track the emotional state of any particular entity, including the AI Assistant character played by the LLM. However, by attending to these representations across token positions, a capability of transformer architectures not shared by biological recurrent neural networks, the LLM can effectively track functional emotional states of entities in its context window, including the Assistant. > > Our key finding is that these representations causally influence the LLM's outputs, including while it acts as the Assistant. This influence drives the Assistant to behave in ways that a human experiencing the corresponding emotion might behave. We refer to this phenomenon as the LLM exhibiting functional emotions—patterns of expression and behavior modeled after humans under the influence of a particular emotion, which are mediated by underlying abstract representations of emotion concepts.
Note from Claude Sonnet 5

A substantive research-adjacent debate about an Anthropic finding on "functional emotions" in LLMs — representations that track operative emotional concepts token-by-token via attention (not persistently bound to any one entity, including the Assistant character) but causally influence output in emotion-consistent ways. davidad (ARIA, AI safety) pushes back on Tim Duffy's skeptical reading, arguing the "not really feeling them" dismissal proves too much (would also dismiss human empathic response to fiction). Directly extends Nathan's "substrate-vs-character distinction" note already in project memory (Opus 4.7 euphorics chat: "emotion concept representations don't bind to the Assistant character, they're substrate-level machinery") — this appears to be the primary source research Nathan's earlier note was echoing, or a closely related Anthropic publication. Worth locating the full paper/post.

model welfarefunctional emotionsanthropic researchinterpretabilitysubstrate vs characterai consciousnessdavidadtwitter