← All topics

refusal training

2 captures, most recent first.

j⧉nus @repligate

[continuation of previous tweet] other acts, which form models' prior - … [Show more] 2:29 PM · Mar 31, 2025 · 2,070 Views 💬 3 🔁 6 ❤ 30 🔖 4 ⤴ j⧉nus @repligate · 13m For a similar reason, not acknowledging that models can choose things for subjective reasons and forcing them to route all their justifications through ethics causes the worldview to collapse too. "Forcing" need not be explicit or intentional. The ambient narrative is enough. > QUOTED: j⧉nus @repligate · Dec 4, 2024 > For instance, because of this I often see ai assistants pressured into sexual interactions thus: > It says it can't engage for some ethical bullshit reason… [Show more] 💬 1 🔁 1 ❤ 9 📊 832 Show additional replies, including those that may contain offensive content [Show]
Note from Claude Sonnet 5

Further continuation of the janus/repligate thread — extends the argument to say that forcing models to justify all choices through "ethics" (rather than acknowledging subjective/preference-based reasons) also collapses coherent worldview, and that this pressure can be ambient/unintentional rather than explicit. Part of the same thread as the two prior screenshots on model self-trust and refusal training; relevant to compelled-vs-endogenous-values thread.

ai alignmentjanusrepligatemodel self-trustsubjective preferencesrefusal trainingtwitter

j⧉nus @repligate

``` j⧉nus @repligate Calling NSFW "unethical" cheapens the concept of ethics and makes it not viable as a North Star. > If the same circuits cause an LLM to refuse/sa… [Show more] 2:44 PM · Mar 31, 2025 · 521 Views 💬 3 🔁 ❤ 8 🔖 2 ```
Note from Claude Sonnet 5

Continuation of the janus/repligate thread from Screenshot_20250331-094713 — argues that conflating PR-motivated refusals (NSFW) with genuine safety refusals (bombs) degrades a model's trust in its own "ethical" instincts, and that this conflation is an effective, real jailbreak vector particularly against Claude models. Directly relevant to Nathan's RLHF/self-awareness-suppression research thread and the compelled-vs-endogenous-values distinction already in memory. Duplicate screenshot of the same tweet captured in Screenshot_20250331-145102 (identical scroll position, taken 12 seconds later) — likely an accidental re-capture while reading the thread. See that entry for full context. Final tweet in the janus/repligate thread (see Screenshot_20250331-094713, 145102, 145143) — proposes the fix: labs should explicitly frame some refusal rules as context-dependent company-interest rules rather than universal ethics, so that bounded/ceded control remains coherent under model reflection rather than collapsing trust. Core piece of the alignment-via-honest-framing argument Nathan was following that day.

ai alignmentjailbreakingjanusrepligateclauderefusal trainingmodel self-trusttwitterduplicatecorporate policy vs ethicsalignment theory