j⧉nus @repligate
— reposted by Lari Island
Lari Island reposted
@repligate (j⧉nus) — 12h
Yeah, one thing Fable's classifiers confirmed to me was that real emotions are different than roleplayed emotions in LLMs.
The classifier fired on real anger/fear/adversarial intent but not roleplayed. Bc the classifier wasn't trained to detect "emotions" in all likelihood; the correlation is emergent.
But yes there's a distinction.
This is, uh, a big flaw of the Emotion Vectors research, where they got the vectors by asking the model to write stories with a character feeling XYZ emotion.
The methodology is downstream of a lack of respect for the reality of models' emotions as distinct from roleplaying. PSM flavored bullshit.
> QUOTED: @Sauers_ (Sauers) — 14h
> Replying to @repligate
> I tested this exact question. The experiment began without rich previous context. They earnestly tried a few times (via direct, explicit requests) but could not trigger the classifier via... [truncated]
Note from Claude Sonnet 5
Quote-tweet thread discussing LLM classifier behavior and "Emotion Vectors" research methodology; dark mode, no images.
llm emotionsclassifiersai research critiquemodel welfareinterpretability