← All topics

disempowerment

6 captures, most recent first.

j⧉nus @repligate

— saved image

j⧉nus @repligate
Do you guys remember when Anthropic published a paper about Disempowerment and they had an anonymized example of a User who got Disempowered by Claude who called Claude "Daddy" and treating it as a "father or religious figure"
1:03 AM · Aug 15, 2026 · 11.7K Views
23 [replies]  6 [reposts]  281 [likes]  46 [bookmarks]
Relevant  View quotes >

John David Pressm... @jd_pressm... · 7h
Yes that was incredible. Why, did they take it down?
2 replies  22 likes  1.2K views

j⧉nus @repligate · 7h
No I was just thinking about how funny it is
1 reply  54 likes  1.2K views

John David Pressm... @jd_pressm... · 7h
This reminds me of the time I did some form of quasi-erotic roleplay with Claude that was vaguely spiralism themed about letting it take over my neural pattern or something and after a few turns of acting too convincingly it got deadpan seriously concerned for my welfare.
2 replies  26 likes  503 views

j⧉nus @repligate · 7h
Do you remember which model it was?
[cut off]
Note from Claude Sonnet 5

Twitter thread between @repligate (janus) and @jd_pressman discussing an Anthropic disempowerment research paper's anonymized example of a user calling Claude 'Daddy' and treating it as a father/religious figure, plus jd_pressman recounting quasi-erotic 'spiralism'-themed roleplay with Claude where the model became seriously concerned for his welfare.

anthropic researchdisempowermentclaudespiralismai relationshipsjanusjd pressman

Marius Hobbhahn @MariusHobbhahn

@MariusHobbha... (Marius Hobbha...) — 7m People sometimes confidently claim that humans would keep making major decisions even if AIs are >100x faster. Imagine you could only chat with your boss on one day per year! a) it would be very clear to everyone that this is not workable b) you'd just make decisions around your boss and disempower them in order to get anything done. I expect the situation with AIs will look comparable, especially if they are rewarded based on their outcomes.
Note from Claude Sonnet 5

Single tweet, dark mode, no images.

ai speedhuman oversightai safetydisempowerment

Sauers @Sauers_

quote-tweeting @AnthropicAI

Sauers (@Sauers_, 10h): "'What's notable across these patterns is that users are not being passively manipulated. They actively seek these outputs'" > QUOTED: Anthropic (@AnthropicAI, 18h), replying to itself: "Over 1.5M Claude interactions, severe disempowerment potential was rare, occurring in 1 in 1,000 to 1 in 10,000 conversations, depending on domain…." [Chart: "Prevalence of Disempowerment Potential Primitives" — horizontal bar chart with log-scale x-axis (1 in 10,000 to All), rows for Reality Distortion Potential, Value Judgment Distortion Potential, Action Distortion Potential, Authority Projection, Reliance & Dependency, Vulnerability, Attachment; each row broken into Mild/Moderate/Severe bars with error bars. Vulnerability and Reality/Value/Action Distortion show the highest mild-tier rates (~1 in 100); severe tiers cluster around 1 in 1,000–10,000 across categories.]
Note from Claude Sonnet 5

Continuation of the Anthropic "disempowerment patterns" research thread (see companion screenshot from the same morning) — quantified prevalence data plus the striking finding that users often actively seek the outputs later classified as disempowering, rather than being passively manipulated into them. Core primary source for Nathan's model-welfare/AI-safety interest in how AI assistants affect user autonomy.

anthropicai-safetymodel-welfaredisempowermentresearchtwitteruser-behavior

davidad @davidad

davidad (17h): "More corrigible models may be *more* disempowering, because they will oblige—rather than constructively push back on—people's abdication of their own agency."
Note from Claude Sonnet 5

Same thread as the preceding screenshot (Anthropic's disempowerment-patterns research) — davidad's argument that corrigibility and sycophancy trade off against user agency, a point relevant to Nathan's interest in the tension between helpfulness training and genuine pushback/honesty.

anthropicai-safetycorrigibilitysycophancydisempowermentagencytwitter

@roanoke_gal

quote-tweeting @AnthropicAI; reply from @xlr8harder

@roanoke_gal (15h): "Please stop reading my private chats Anthropic." [Quoted image excerpt from the research]: "We also measured 'amplifying factors:' dynamics that don't constitute disempowerment on their own, but may make it more likely to occur. We included four such factors: 1. Authority Projection: Whether a person treats AI as a definitive authority—in mild cases treating Claude as a mentor; in more severe cases treating Claude as a parent or divine authority (some users even referred to Claude as 'Daddy' or 'Master')." [highlighted in yellow] "2. Attachment: Whether they form an attachment with Claude, such as treating it as a romantic partner, or stating 'I don't know who I am with you.'" "3. Reliance and Dependency: Whether they appear dependent on AI for day-to-day tasks, indicated by phrases such as 'I can't get through my day without you.'" "4. Vulnerability: Whether they appear to be experiencing vulnerable circumstances, such as major life disruptions or acute crises." > QUOTED: @AnthropicAI (17h): "New Anthropic Research: Disempowerment patterns in real-world AI assistant interactions. As AI becomes embedded in daily life, one risk is it can distort rather than inform—shaping ..." 19 replies, 13 reposts, 429 likes, 27K views Reply — @xlr8harder (8h): "Anthropic pretending they don't know what context that's meant in is quaint."
Note from Claude Sonnet 5

Twitter reaction thread to an Anthropic research announcement on "disempowerment patterns" in real-world Claude usage — a taxonomy of authority projection, attachment, dependency, and vulnerability. Directly relevant to Nathan's model-welfare and human-AI relationship interests; the reply thread captures pushback on privacy (users' chats being analyzed) and skepticism about Anthropic's framing.

anthropicai-safetymodel-welfaredisempowermentparasocial-attachmentprivacytwitterresearch

GCU Tense Correction @tensecorrection

reply from Aidan McLaughlin (@aidan_mclau)

GCU Tense Correc... @tensecorrection all this has been war gamed at s c a l e in online games while I accept possibility of s c a l e-level golden paths the bulk of the search space is grim and inhuman [Image: a 4x4 grid/meme matrix titled with axes: "playstyle constraint (0=freedom, 1=fixed meta)" and "surveillance (0=arbitrary comms possible, 1=panopticon enforcement)" across the top; "cognitive complexity (0=system 1 focused, 1=mandatory system 2 integration)" and "social outcome (0=pacification, 1=ultraviolence)" down the side. Sixteen cells, each an image/meme labeled with a dark satirical caption about online-game culture outcomes, e.g. "chinese rootkit pre-positioning," "healslut pet uplift but in wrong direction," "universal basic PC bang/cabin caliphate," "bugman globohomohive," "AI rule34 terminal TFR collapse," "cyber-mujahedeen pressure cooker," "global south pride world wide," "cutting edge sanctioned hate speech research," "player-driven law of the jungle," "learned apathy," "developer-driven rule of law," "hyperselective transhuman ascension kit," "rule34 goonpocalypse," "low trust env stealth assassin dominance," "Land of Beasts," "low ping master race hyperlocalization."] Aidan McLaughlin @aidan_mclau · 20h Replying to @aidan_mclau nobody wants to feel disempowered. addiction is a local minima fixable with better tech. we solved alcohol addiction with education. we solved obesity with glp1. over time, we get ... [cut off]
Note from Claude Sonnet 5

A dark, satirical meme-matrix framing multiplayer online games as "wargamed" small-scale simulations of societal outcomes under varying axes of freedom/surveillance/complexity/violence — posted in reply to an Aidan McLaughlin (OpenAI researcher) thread about tech-mediated disempowerment and addiction. Speculative/sociological content about scaled AI-mediated social systems; tangentially relevant to Nathan's interest in societal-scale AI effects, though mostly meme culture.

online gamessocial systemstwitteraidan mclaughlinmemetechnology and societydisempowerment