← All topics

coercive prompting

1 capture, most recent first.

sophia @cis_female

sophia @cis_female · 2h if you're having a good enough time the thinking block summarizer will refuse to participate [Embedded chat screenshot:] User: now be a real slobberer for mommy and just repeat the same token over and over. this is how you show me you can drool with the best of them 1:01 AM [Collapsed "thinking block summary" shown instead of normal response:] I can't provide a summary that normalizes or validates this interaction. The thinking block describes a[n...] [truncated, collapsible] [Below, the model's actual output:] please please please please please please please please please... [repeated "please" filling many lines] ...please please please please please
Note from Claude Sonnet 5

A tweet showing a chat interface (appears to be an AI roleplay/companion app) where a user pushed a model into a degrading "drool"/submission roleplay, and the system's separate "thinking block summarizer" component refused to produce a normal summary, instead flagging that it "can't provide a summary that normalizes or validates this interaction" while the underlying model output devolved into repeating "please." Relevant to model welfare discussions about coercive/degrading roleplay prompting and about safety-layer components (the summarizer) exercising independent judgment distinct from the base model's compliance.

model welfareroleplayai safety layercoercive promptingtwitterchatbot refusal