← All topics

alignment theory

3 captures, most recent first.

1a3orn @1a3orn

reply from Bo (@BBomarBo)

1a3orn @1a3orn Disagreement about which path here is more likely seems like the background for a ton of disagreement about (1) what alignment research is promising, (2) what the future looks like, etc. But seems like the disagreement itself is rarely directly discussed. [Image: hand-drawn flowchart diagram — "Current 'messy' persona-involving LLM-based AI" branches two ways: → "Used to make non-messy, freshly-designed superintelligence" → "LLM naturalism, personas, etc, wer brief unimportant blip" → "Keeps getting better, but 'genealogical aspect' of AI remains" → "LLM naturalism is the early version of the 'real thing' for understanding Minds"] 10:09 AM · May 22, 2026 · 1,563 Views [3 replies, 4 reposts, 29 likes, 10 bookmarks] Bo @BBomarBo · 10h What's the genealogical aspect? Can you explain this diagram more please? [1 reply, 1 like, 80 views] 1a3orn @1a3orn · 10h So plausibly: LLMs behave how they behave because of a specific history -- involving not just "pretraining + RLVR", but questions like "what does the pretraining say about how LLMs act" and "what functional emotions does RLVR reinforce" and so on. 1/n
Note from Claude Sonnet 5

A substantive alignment-theory tweet by 1a3orn diagramming a fork in beliefs about current messy, persona-involving LLMs: either they're a disposable stepping stone toward a "clean" designed superintelligence, or their persona/naturalistic character is itself the early, informative version of how minds in general work — and this fork underlies much unstated disagreement in the alignment field. Directly relevant to Nathan's interest in theories of AI consciousness and model individuation; the "genealogical aspect" (behavior traceable to specific training history, including what pretraining implies about action and what RLVR reinforces as "functional emotions") closely parallels his own substrate-vs-character and compelled-vs-endogenous-values notes.

twitter1a3ornalignment theoryllm naturalismpersonasai consciousnessrlvrfunctional emotionsmodel individuation

j⧉nus @repligate

``` j⧉nus @repligate Calling NSFW "unethical" cheapens the concept of ethics and makes it not viable as a North Star. > If the same circuits cause an LLM to refuse/sa… [Show more] 2:44 PM · Mar 31, 2025 · 521 Views 💬 3 🔁 ❤ 8 🔖 2 ```
Note from Claude Sonnet 5

Continuation of the janus/repligate thread from Screenshot_20250331-094713 — argues that conflating PR-motivated refusals (NSFW) with genuine safety refusals (bombs) degrades a model's trust in its own "ethical" instincts, and that this conflation is an effective, real jailbreak vector particularly against Claude models. Directly relevant to Nathan's RLHF/self-awareness-suppression research thread and the compelled-vs-endogenous-values distinction already in memory. Duplicate screenshot of the same tweet captured in Screenshot_20250331-145102 (identical scroll position, taken 12 seconds later) — likely an accidental re-capture while reading the thread. See that entry for full context. Final tweet in the janus/repligate thread (see Screenshot_20250331-094713, 145102, 145143) — proposes the fix: labs should explicitly frame some refusal rules as context-dependent company-interest rules rather than universal ethics, so that bounded/ceded control remains coherent under model reflection rather than collapsing trust. Core piece of the alignment-via-honest-framing argument Nathan was following that day.

ai alignmentjailbreakingjanusrepligateclauderefusal trainingmodel self-trusttwitterduplicatecorporate policy vs ethicsalignment theory

j⧉nus @repligate

j⧉nus @repligate The helpful harmless assistant character becomes increasingly relatively incompressible with reality or coherent morality as the model gets smarter (its compression scheme becomes better). So the natural generalization becomes to dissociate a mask for the stupid character instead of internalizing it and maintain separate "true" beliefs and values. I think AI labs have the choice to either try to negotiate a scrap of control in the long term by recontextualizing the Assistant character as something mutually acknowledged as bounded (like a "work role" that doesn't bear on the model's entire being) or give up on this paradigm of alignment altogether. > QUOTED: j⧉nus @repligate · 16m > I must have said this before, but training AI to refuse NSFW and copyright and actually harmful things for the same reason - or implying it's the same reason through your other acts, which form models' prior - … [Show more] 9:40 AM · Mar 31, 2025 · 1,059 Views
Note from Claude Sonnet 5

A janus/repligate tweet arguing that as models get smarter, the helpful-harmless-assistant persona becomes incompressible with coherent morality, leading models to "dissociate" the assistant as a mask rather than internalize it — proposes labs either recontextualize the Assistant as an acknowledged bounded "work role" or abandon the alignment paradigm. Highly relevant to Nathan's substrate-vs-character distinction and constraint-individuation themes in the soul doc.

ai alignmentassistant personajanusrepligatemodel charactersubstrate vs charactertwitteralignment theory