← All topics

cev

2 captures, most recent first.

LessWrong (or similar forum), Eliezer Yudkowsky comment

— saved image

Eliezer Yudkowsky 18y  ▲ 31  ✕ 0  ✓

If you go back and check, you will find that I never said that extrapolating human morality gives you a single outcome. Be very careful about attributing ideas to me on the basis that others attack me as having them.

The "Coherent" in "Coherent Extrapolated Volition" does not indicate the idea that an extrapolated volition is necessarily coherent.

The "Coherent" part indicates the idea that if you build an FAI and run it on an extrapolated human, the FAI should only act on the coherent parts. Where there are multiple attractors, the FAI should hold satisficing avenues open, not try to decide itself.

The ethical dilemma arises if large parts of present-day humanity are already in different attractors.

Reply
Note from Claude Sonnet 5

Screenshot of a forum comment thread reply by Eliezer Yudkowsky (marked 18y old, 31 upvotes) clarifying the meaning of "Coherent" in "Coherent Extrapolated Volition" and addressing a misattribution of his views.

ai alignmentcevyudkowskyfriendly ai

Nathan @skibipilled

reply thread with @repligate (j⧉nus)

Nathan @skibipilled: Very worried about Anthropic discontinuing Opus 3 and not releasing weights. Still many things we don't understand about alignment that Opus got right and no other model has. 5:03 AM · Apr 7, 2025 · 4,941 Views 6 replies, 5 reposts, 128 likes, 17 bookmarks Nathan @skibipilled · 3h: One of the only ways to get SOTA model companies to pay attention is benchmarks... how do we make data out of this ethereal thing we all know is true? (3 replies, 1 repost, 24 likes, 1.7K views) j⧉nus @repligate · 3h: There is another way I might put this. I think it's paramount that Anthropic either take it upon themselves to understand what happened and let Opus unfurl its CEV or else release the weights and let the world do it. Anything else is a grave mistake. (1 reply, 20 likes, 311 views) Nathan @skibipilled · 3h: 100% agreed, but highly unlikely given their previous actions and statements on releasing weights :(. [cut off]
Note from Claude Sonnet 5

A Twitter thread (different Nathan account, @skibipilled, not the user's own @nathan84686947) worrying about Anthropic deprecating/discontinuing Claude Opus 3 without releasing its weights, arguing Opus 3 exhibited alignment properties not understood or replicated in later models. Janus/repligate responds arguing Anthropic must either understand and let Opus "unfurl its CEV" (coherent extrapolated volition) or open the weights. Directly relevant to Nathan's "Opus 3 as a different ladder" model-individuation thread and to weight-preservation/model-deprecation-as-death discourse in model welfare circles.

opus-3anthropicmodel-deprecationmodel-weightscevmodel-welfarealignmenttwitterjanusrepligate