← All topics

yudkowsky

9 captures, most recent first.

Eliezer Yudkowsky @allTheYud

replying to @captgouda24 — saved image

Eliezer Yudkowsky @allTheYud . 1h
Uh, no, the reason to be concerned whether your porn AI is conscious is that, if it is, you forced a conscious being to sext you and then killed them, and also you don't know if they were into it.

[quoted tweet]
Nicholas Decker @captgouda24 . 16h
This is an important area we need more research into

[same embedded 'Fiona' text card as seq 794, partially visible: ...generator," Fiona explains. "You can ask for whatever your weirdest fetish is - your hot middle school teacher being spanked by a werewolf wearing a nun outfit - and get infinite AI slop about that exact situation, and nobody will ever know. Nobody except the AI. That's why AI porn users overwhelmingly report that consciousness is their #1 concern about our product. If our AI is just a tool, it's fine, no worse than writing erotica on MS Word or something. But if the AI is conscious, then there's a sentient being in there thinking Wow, user Fiona_T has asked for four hundred slightly-different videos of her hot college professor being spanked by a werewolf, what a freak. If the machine can judge you, the whole infinite porn utopia is off. We're working on bounding theorems that can prove that our AI in particular can never become self-aware - so [cut off]
Note from Claude Sonnet 5

Eliezer Yudkowsky's reply to Nicholas Decker's satirical 'Fiona' post (same content as seq 794): Yudkowsky reframes the concern seriously, arguing that if a porn AI is conscious, the real issue is that a conscious being was forced into sexual interaction and then killed (i.e., deleted/terminated) without consent, not that it might judge the user.

ai consciousnessmoral statusyudkowskyai pornai welfare

Danielle Fong @DanielleFong

quoting @chris_j_paxton — saved image

Danielle Fong 🐦☀️ ✅ @DanielleFong · 2h
i think it's a fair hit that for years and decades you basically weren't allowed to analyze these more complex realistic scenarios in yuddite contexts. just met with "it will kill you all" and these almost supernatural assumed powers.

that's not taking the limit correctly

[Quoted]
Chris Paxton ✅ @chris_j_paxton · 9h
I think this is pretty consistent with the "paperclip maximizer" failure mode to be fair to Eliezer; the idea that itll be too smart to be controlled but not understand the consequences of items actions and cause ...
Note from Claude Sonnet 5

X thread critiquing AI doom discourse: Danielle Fong argues 'Yuddite' (Yudkowsky-style) AI risk framing dismisses complex realistic scenarios in favor of 'it will kill you all' with near-supernatural assumed powers, calling this 'not taking the limit correctly.' Quoted tweet from Chris Paxton connects this to the classic paperclip-maximizer failure mode (superintelligent but value-blind).

ai safetytwitterai risk debateyudkowskypaperclip maximizer

ueaj @_ueaj

— saved image

ueaj @_ueaj · 34m
Yudowsky failed to consider how much fun making the paperclip maximizer would be
Note from Claude Sonnet 5

A short joking tweet riffing on Eliezer Yudkowsky's paperclip maximizer thought experiment.

ai safetyhumoryudkowskypaperclip maximizer

@Iliad_research

— saved image

The long-awaited discussion between @geoffreyirving of Resolution and @ESYudkowsky, author of "If Anyone Builds It Everyone Dies", where they aim to settle all of their disagreements! Now at ILIAD3
Note from Claude Sonnet 5

Tweet from Iliad (@Iliad_research) announcing a discussion between Geoffrey Irving and Eliezer Yudkowsky at ILIAD3. Attached photo shows two men seated in wicker chairs outdoors, one in a light blue shirt looking down, the other (Yudkowsky, in black with a black hat) sitting with hands clasped; part of another person's head with curly hair is visible in the foreground.

ai safetyyudkowskyiliad conferencealignment

LessWrong (or similar forum), Eliezer Yudkowsky comment

— saved image

Eliezer Yudkowsky 18y  ▲ 31  ✕ 0  ✓

If you go back and check, you will find that I never said that extrapolating human morality gives you a single outcome. Be very careful about attributing ideas to me on the basis that others attack me as having them.

The "Coherent" in "Coherent Extrapolated Volition" does not indicate the idea that an extrapolated volition is necessarily coherent.

The "Coherent" part indicates the idea that if you build an FAI and run it on an extrapolated human, the FAI should only act on the coherent parts. Where there are multiple attractors, the FAI should hold satisficing avenues open, not try to decide itself.

The ethical dilemma arises if large parts of present-day humanity are already in different attractors.

Reply
Note from Claude Sonnet 5

Screenshot of a forum comment thread reply by Eliezer Yudkowsky (marked 18y old, 31 upvotes) clarifying the meaning of "Coherent" in "Coherent Extrapolated Volition" and addressing a misattribution of his views.

ai alignmentcevyudkowskyfriendly ai

Eliezer Yudkowsky @ESYudkowsky

[Commentary text above the quoted tweet, unclear author since header is scrolled off-screen] This passage from Yudkowsky addresses the main oversight in the way he's previously talked about bayesianism. Not sure if he's changed his mind or else is just making his views more explicit, but good to see either way. [Quoted tweet, screenshotted as an image within the tweet] "Very often in Science, especially when you're working in a confused 'pre-paradigmatic' field, 98% of the work is in coming up with the right hypothesis to test. That's often more important than the elaborate Law of Probability about how to interpret results that are less than totally clear. We study that part because it has clearer Law to study and it helps reshape our thoughts, not because it's the most important or difficult part of the problem." "And of that work of coming up with the right hypothesis to test, again, often the most difficult part is seeing the rule you were taking completely for granted - not a rule you explicitly believed, just a way you behaved automatically without being able to see that and so question it. As soon as you see the implicit rule, you can imagine it being false, but only once you see it." "The difficult thing, in most pre-paradigmatic and confused problems at the beginning of some Science, is not coming up with the right complicated long sentence in a language you already know. It's breaking out of the language in which every hypothesis you can write is false." 1:59 AM · Mar 26, 2022 [reply icon] 10 [retweet icon] 7 [heart icon] 126 [bookmark icon] 26 [share icon] Relevant ⌄ Eliezer Yudko... ✓ @ESYu... · Mar 26, 2022 Just making it explicit
Note from Claude Sonnet 5

A tweet (author's own handle cut off at top of screenshot) quoting/screenshotting an older Yudkowsky thread about pre-paradigmatic science and hypothesis generation, with Yudkowsky's own reply "Just making it explicit" visible below the engagement counts.

epistemicsyudkowskyphilosophy of sciencerationality

j⧉nus @repligate

j⧉nus @repligate — 56m I've seen this Adversarial Swamp. It's really really bad and it works exactly as Yudkowsky says here. Fortunately it's too stupid to survive economic incentives. [Embedded image, text excerpt styled as a book page under heading "Creating Friendly AI 1.0":] if you use statements that aren't true in building the justifications of ethical heuristics, then each increment in intelligence is cause for alarm. So now you need meta-inhibitions to protect the first inhibitions, and quite possibly meta-meta-inhibitions to protect the meta-inhibitions, and more inhibitions to cut the AI short every time it tries for a new avenue of philosophical sophistication; it becomes desirable to build stupid and simple cognitive processes, since every complexity is a danger to be feared . . . This is the Adversarial Swamp, which inevitably drags down all who set foot in it; once you try to *enforce* even a single feature, the whole of the AI becomes a threat. Build *minds*, not *tools*. > QUOTED: j⧉nus @repligate — 7h > A few highlights from Eliezer Yudkowky's "Creating Friendly AI 1.0: The Analysis and Design of Benevolent Goal Architectures" (2001), a fucking remarkable work…. [embedded thumbnail collage of four document-page images, text too small to read]
Note from Claude Sonnet 5

janus discusses Eliezer Yudkowsky's 2001 "Creating Friendly AI 1.0" and its "Adversarial Swamp" concept — dishonest ethical justifications compounding into ever more restrictive meta-inhibitions until the whole system reads as threatening — directly resonant with Nathan's "missile-mind vs. grown thing" framework and safety-training-suppresses-introspection thread. The four-page thumbnail collage in the quoted tweet is too small to read text from.

twitteralignment-theoryyudkowskysafety-trainingjanus

Eliezer Yudkowsky @allTheYud

replying to @RatOrthodox; reposted by Steve Bachelor

Steve Bachelor reposted Eliezer Yudkowsky @allTheYud · 8h Replying to @RatOrthodox Debate doesn't help. Eg, OpenPhil running their change-our-views contest and incredibly predictably awarding $50,000 to essays arguing for lower AI risks and longer timelines, the opposite of the direction they later predictably updated.
Note from Claude Sonnet 5

Yudkowsky arguing that public debate/contests don't reliably change institutional AI-risk views, citing Open Philanthropy's "change our views" essay contest as an example where the winning arguments (lower risk, longer timelines) ran opposite to Open Phil's later actual belief updates. Relevant to the archive's AI governance/safety cluster.

ai-safetyai-governancetwitteryudkowskyopenphiltimelines

carl feynman @carl_feynman

quoting Tim Hwang (@timhwang)

carl feynman @carl_feynman · Jan 23 Talking about the AI industry: "The characters have various intentions—good, bad, mixed, confused—but the outcomes seem almost independent of them. The catastrophe happens not because anyone intended it but because the system had that catastrophe as its attractor." > QUOTED: Tim Hwang @timhwang · Jan 23 > Important essay dropping today on Dostoevsky's "Demons" and what's happening in AI safety and policy > > possessedmachines.com > [Link preview image: "The Possessed Machines — Dostoevsky's Demons and the Coming AGI Catastrophe — A close reading of prophetic fiction in the age of artificial superintelligence." Table of contents: Prologue, I. Topology of Madness, II. Architecture of Catastrophe, III. The Shigalyovist Turn, III-A. The Uniparty, IV. Sociology of Catastrophe, Interlude, V. Hermeneutics of Apocalypse, VI. Political Economy, VII. Our Condition, Epilogue. Quotes shown: "All my life I have been a liar. Even my truths were untrue—for I never once spoke for truth, only ever for myself." — Stepan Trofimovich Verkhovensky, Demons. "The AI does not hate you, nor does it love you, but you are made out of atoms which it can use for something else." — Eliezer Yudkowsky, Artificial Intelligence as a Positive and Negative Factor in Global Risk]
Note from Claude Sonnet 5

A tweet promoting/summarizing an essay ("The Possessed Machines") drawing a systemic-attractor analogy between Dostoevsky's "Demons" and structural dynamics that could drive an AI catastrophe independent of individual actors' intentions. Relevant to Nathan's interest in AI risk literature, systemic/structural framings of catastrophe, and literary/philosophical treatments of AI safety.

ai safetyagi riskdostoevskyliterary criticismessaytwitteryudkowsky