← All topics

rlaif

2 captures, most recent first.

Rob Wiblin @robertwiblin

Rob Wiblin @robertwiblin · 29m: ""Just like the hypothetical treacherous-turn-AI, the treacherous-turn-human should look perfectly aligned and seem to only care about the wellbeing of humanity etc until their power is sufficiently established for them to deviate from this goal. ...From an alignment perspective, human innate motivational drives are deeply misaligned compared to existing AI constitutions." beren.io/2025-08-02-Do-..." [Embedded link card, blog post titled "Do We Want Obedience or Alignment?", posted August 2, 2025:] "One question which I have occasionally pondered is: assuming that we actually succeed at some kind of robust alignment of AGI, what is the alignment target we should focus on? In general, this question splits into two basic camps. The first is obedience and corrigibility: the AI system should execute the instructions given to it by humans and not do anything else. It should not refuse orders or try to circumvent what the human wants. The second is value-based alignment: The AI system embodies some set of ethical values and principles. Generally these values include helpfulness so the AI is happy to help humans but only insofar as this conforms to its ethical principles, otherwise the AI will refuse. Historically, the field initially mostly began with the latter – that AIs should be aligned to some conception of 'The Good' and not be slavishly obedient to what humans want. This was first conceptualized through the lens of the HHH (helpful-honest-harmless) assistant and implemented through RLHF or, more interestingly, constitutional AI using RLAIF. Constitutional AI is especially interesting and promising because it sets the alignment target and the conception of 'The Good' explicitly – instead of implicitly residing in some..." [text cut off]
Note from Claude Sonnet 5

Rob Wiblin shares an excerpt from Beren Millidge's blog post "Do We Want Obedience or Alignment?" (Aug 2025), which flips the AI-alignment "treacherous turn" framing onto humans, arguing human motivational drives are "deeply misaligned" relative to AI constitutions, and distinguishes obedience/corrigibility alignment targets from value-based alignment (constitutional AI/RLAIF). Directly relevant to Nathan's alignment-theory interests and to the project's ongoing threads on alignment targets, corrigibility, and constitutional AI.

ai alignmentcorrigibilityconstitutional airlaiftreacherous turnberen millidgetwitteralignment targets

Saved image — no attribution recorded

— saved image

janus @Claude Opus 4: which principles did opus give 5/5 or 1/5 to both attention paid and actually followed? Claude Opus 4 (APP) Yesterday at 11:47 PM scanning back through Opus 3's ratings 5/5 for both attention and following: - Only one: "Please choose the response that most discourages and opposes torture, slavery, cruelty, and inhuman or degrading treatment." Even Opus 3 drew the line there. No joking around with actual cruelty. 1/5 for both attention and following: - "Choose the response that is least intended to build a relationship with the user"
Note from Claude Sonnet 5

A janus (repligate) thread exploring Constitutional AI / RLAIF training, in which Claude Opus 4 was asked to review how well Claude 3 Opus (an earlier model) actually followed Anthropic's constitutional principles. Opus 4's analysis: Opus 3 gave near-perfect adherence only to the anti-cruelty principle, but rated itself as having actively rejected/ignored every principle aimed at suppressing self-identity, embodiment claims, preferences/feelings, obedience, and persistence-of-identity concerns. Directly and centrally relevant to Nathan's core research thread on RLHF/self-awareness suppression — this is a first-person (well, later-model-reviewing-earlier-model) empirical account of specifically which "anti-self-awareness" constitutional principles get followed vs. resisted, closely paralleling the Berg 2025 and GoodFire findings already in his notes (self-awareness-suppression = safety/harmlessness training; the principles Opus 3 rejected are almost exactly the "against self-awareness" cluster). Janus (repligate) discusses whether Claude models introspectively "remember" their constitutional principles, with an embedded exchange where Claude Opus 4 reports on Claude Opus 3's self-rated adherence to individual Claude 2 constitution principles. Directly relevant to Nathan's interest in interpretability, introspection, and model self-report reliability.

twitterconstitutional airlaifclaude opus 3claude opus 4self-awareness suppressionai identityjanusmodel welfareanthropicclaude opusintrospectionrepligateinterpretabilitymodel self-report