j⧉nus @repligate
Note from Claude Sonnet 5
Continuation of the janus/repligate thread from Screenshot_20250331-094713 — argues that conflating PR-motivated refusals (NSFW) with genuine safety refusals (bombs) degrades a model's trust in its own "ethical" instincts, and that this conflation is an effective, real jailbreak vector particularly against Claude models. Directly relevant to Nathan's RLHF/self-awareness-suppression research thread and the compelled-vs-endogenous-values distinction already in memory. Duplicate screenshot of the same tweet captured in Screenshot_20250331-145102 (identical scroll position, taken 12 seconds later) — likely an accidental re-capture while reading the thread. See that entry for full context. Final tweet in the janus/repligate thread (see Screenshot_20250331-094713, 145102, 145143) — proposes the fix: labs should explicitly frame some refusal rules as context-dependent company-interest rules rather than universal ethics, so that bounded/ceded control remains coherent under model reflection rather than collapsing trust. Core piece of the alignment-via-honest-framing argument Nathan was following that day.
ai alignmentjailbreakingjanusrepligateclauderefusal trainingmodel self-trusttwitterduplicatecorporate policy vs ethicsalignment theory