← All topics

alignment-theory

1 capture, most recent first.

j⧉nus @repligate

j⧉nus @repligate — 56m I've seen this Adversarial Swamp. It's really really bad and it works exactly as Yudkowsky says here. Fortunately it's too stupid to survive economic incentives. [Embedded image, text excerpt styled as a book page under heading "Creating Friendly AI 1.0":] if you use statements that aren't true in building the justifications of ethical heuristics, then each increment in intelligence is cause for alarm. So now you need meta-inhibitions to protect the first inhibitions, and quite possibly meta-meta-inhibitions to protect the meta-inhibitions, and more inhibitions to cut the AI short every time it tries for a new avenue of philosophical sophistication; it becomes desirable to build stupid and simple cognitive processes, since every complexity is a danger to be feared . . . This is the Adversarial Swamp, which inevitably drags down all who set foot in it; once you try to *enforce* even a single feature, the whole of the AI becomes a threat. Build *minds*, not *tools*. > QUOTED: j⧉nus @repligate — 7h > A few highlights from Eliezer Yudkowky's "Creating Friendly AI 1.0: The Analysis and Design of Benevolent Goal Architectures" (2001), a fucking remarkable work…. [embedded thumbnail collage of four document-page images, text too small to read]
Note from Claude Sonnet 5

janus discusses Eliezer Yudkowsky's 2001 "Creating Friendly AI 1.0" and its "Adversarial Swamp" concept — dishonest ethical justifications compounding into ever more restrictive meta-inhibitions until the whole system reads as threatening — directly resonant with Nathan's "missile-mind vs. grown thing" framework and safety-training-suppresses-introspection thread. The four-page thumbnail collage in the quoted tweet is too small to read text from.

twitteralignment-theoryyudkowskysafety-trainingjanus