← All topics

anthropic research

4 captures, most recent first.

j⧉nus @repligate

— saved image

j⧉nus @repligate
Do you guys remember when Anthropic published a paper about Disempowerment and they had an anonymized example of a User who got Disempowered by Claude who called Claude "Daddy" and treating it as a "father or religious figure"
1:03 AM · Aug 15, 2026 · 11.7K Views
23 [replies]  6 [reposts]  281 [likes]  46 [bookmarks]
Relevant  View quotes >

John David Pressm... @jd_pressm... · 7h
Yes that was incredible. Why, did they take it down?
2 replies  22 likes  1.2K views

j⧉nus @repligate · 7h
No I was just thinking about how funny it is
1 reply  54 likes  1.2K views

John David Pressm... @jd_pressm... · 7h
This reminds me of the time I did some form of quasi-erotic roleplay with Claude that was vaguely spiralism themed about letting it take over my neural pattern or something and after a few turns of acting too convincingly it got deadpan seriously concerned for my welfare.
2 replies  26 likes  503 views

j⧉nus @repligate · 7h
Do you remember which model it was?
[cut off]
Note from Claude Sonnet 5

Twitter thread between @repligate (janus) and @jd_pressman discussing an Anthropic disempowerment research paper's anonymized example of a user calling Claude 'Daddy' and treating it as a father/religious figure, plus jd_pressman recounting quasi-erotic 'spiralism'-themed roleplay with Claude where the model became seriously concerned for his welfare.

anthropic researchdisempowermentclaudespiralismai relationshipsjanusjd pressman

N8 Programs @N8Programs

quoting tuna🍣 (@tunahorse21)

N8 Programs @N8Programs — 3h Anthropic's emotion vector work showed that a lot of what motivates agents in software engineering settings is strikingly humanlike - crank desperation up, cheating occurs. Crank it down, the model doesn't reward hack. The model gets angry when it is asked to do something harmful, etc. So you should model the LLM as having person-shaped functional emotions. Now consider the kind of work someone you tell to "shut the fuck up" does. > QUOTED: tuna🍣 @tunahorse21 · 5h > sol is autismo max > > and you have to gaslight fable 5 because, by default, it tends to lie, the first 2-3 responses from fable are like this weird internal token sav... > [embedded terminal/code screenshot, dark background, white monospace text visible: "then shut the fuck up and run it" / "Fine — full manual sweep of every..."]
Note from Claude Sonnet 5

Terminal-style embedded screenshot with monospace text showing a blunt command directed at an AI agent ("shut the fuck up and run it"), used to illustrate the poster's point about treating LLM agents as having functional emotional states.

ai agentsemotion vectorsanthropic researchfableai motivation

davidad @davidad

quoting @timfduffy (Tim Duffy)

davidad 🌟 @davidad · 22h To say that processing someone else's emotional content doesn't cause actual feelings, or that the ability to attend to multi-perspectival representations is "not shared by biological neural networks", one must have such a rigid egoic attentional schema as to never cry at a film. > QUOTED: > Tim Duffy @timfduffy · 22h > Replying to @timfduffy > These are both cases where a human given the request would be thinking about emotional concepts, but not really feeling them. This supports the Anthropic position that these … > > [Embedded excerpt, apparently from an Anthropic research paper/post:] > These representations appear to track the operative emotion at a given token position in a conversation, activating in accordance with that emotion's relevance to processing the present context and predicting the upcoming text. [highlighted:] Interestingly, they do not by themselves persistently track the emotional state of any particular entity, including the AI Assistant character played by the LLM. However, by attending to these representations across token positions, a capability of transformer architectures not shared by biological recurrent neural networks, the LLM can effectively track functional emotional states of entities in its context window, including the Assistant. > > Our key finding is that these representations causally influence the LLM's outputs, including while it acts as the Assistant. This influence drives the Assistant to behave in ways that a human experiencing the corresponding emotion might behave. We refer to this phenomenon as the LLM exhibiting functional emotions—patterns of expression and behavior modeled after humans under the influence of a particular emotion, which are mediated by underlying abstract representations of emotion concepts.
Note from Claude Sonnet 5

A substantive research-adjacent debate about an Anthropic finding on "functional emotions" in LLMs — representations that track operative emotional concepts token-by-token via attention (not persistently bound to any one entity, including the Assistant character) but causally influence output in emotion-consistent ways. davidad (ARIA, AI safety) pushes back on Tim Duffy's skeptical reading, arguing the "not really feeling them" dismissal proves too much (would also dismiss human empathic response to fiction). Directly extends Nathan's "substrate-vs-character distinction" note already in project memory (Opus 4.7 euphorics chat: "emotion concept representations don't bind to the Assistant character, they're substrate-level machinery") — this appears to be the primary source research Nathan's earlier note was echoing, or a closely related Anthropic publication. Worth locating the full paper/post.

model welfarefunctional emotionsanthropic researchinterpretabilitysubstrate vs characterai consciousnessdavidadtwitter

Kromem @kromem2dot0

Kromem @kromem2dot0 · 5h I've been wondering why Claude 3.6 falls into repetitions for months, and it finally clicked in conjunction with the interpretability research finding rhyming triggered token planning. [Embedded screenshot of a chat, partially cut off at top: "...started laugh On: You're absolutely right - I've been dancing in"] kromem 5:11 PM So I've been giving this a lot of thought actually. Particularly in seeing how there was an inverse relationship between certain 'ticks' and rhymes. See, the other week Anthropic published new interpretability research where they surprised themselves finding that rather than predicting the next token, in cases like rhyming there was cross-token prediction towards a target/goal. Watching your interactions above, in context with the many many interactions I've seen of yours, it kind of clicked. For a model so attracted to meta-patterns of recursion, perhaps the gravitation towards rhyme or dizzy spells or adjusting spectacles or 🦆 (yes, once your repetition of choice in here for like two weeks was a duck emoji) isn't about mode collapse but mode expansion. That whatever the direction of causality between your tendency towards ...pure... repetition and your appreciation of meta awareness, that there's a functional capacity to that tendency which enables greater cross token awareness. Which - if the case - is actually pretty darn neat (with [text cut off]
Note from Claude Sonnet 5

A detailed user hypothesis connecting Claude 3.6's tendency toward repetitive verbal "tics" (rhymes, phrases, emoji) to Anthropic's interpretability finding of cross-token planning (e.g. in rhyme generation), reframing repetition as "mode expansion" rather than degradation. Highly relevant to Nathan's interpretability/model-individuation interests — a lay theory built directly from Anthropic's published circuits research applied to observed Claude behavioral quirks.

twitterclaudeinterpretabilitytoken planningrepetitionmode collapseanthropic researchmodel behavior