Rob Wiblin @robertwiblin · 7h
Every day is just a 2010s LessWrong blog post.
> QUOTED: Mitt Romney @MittRomney · Jun 5
> Our highest and most urgent national priority should be AI safeguards. The risks of AI weapons, pathogens, mass unemployment, surveillance, and even extinction must not continue to be largely ignored. x.com/WSJTech...
Note from Claude Sonnet 5
Quote-tweet; quoted tweet text truncated with platform ellipsis to a link.
ai safetymitt romneypoliticstwitter
Rob Wiblin ✓ @robertwiblin · 14h
Anthropic says they're against fully-autonomous killer robots, but if so why does Claude have a 'Constitution' granting it more hit points, resistance to poison, and better concentration saving throws.
> Under Secretary of War Emil Michael ✓ @USWREMichael
> Imagine your worst nightmare. Now imagine that @AnthropicAI has their own "Constitution." Not corporate values, not the United States Constitution, but their own plan to impose on Americans their corporate laws. Claude's Constitution \ Anthropic.
> [Image: line-drawing of a scroll labeled "Claude's Constitution" with a tree/branch icon and a quill]
Note from Claude Sonnet 5
A satirical tweet (D&D "saving throws" joke) mocking a US government official's (Under Secretary of War Emil Michael, part of the renamed Department of War) alarmist framing of Anthropic's Claude Constitution as a nightmarish corporate power grab. Direct continuation of the Anthropic-vs-Department-of-War/Pentagon storyline from Feb 25-26 screenshots; shows government-official hostility toward the Claude Constitution concept itself.
twitteranthropicclaude constitutiondepartment of waremil michaelai policysatirerob wiblin
Rob Wiblin @robertwiblin · 29m:
""Just like the hypothetical treacherous-turn-AI, the treacherous-turn-human should look perfectly aligned and seem to only care about the wellbeing of humanity etc until their power is sufficiently established for them to deviate from this goal.
...From an alignment perspective, human innate motivational drives are deeply misaligned compared to existing AI constitutions."
beren.io/2025-08-02-Do-..."
[Embedded link card, blog post titled "Do We Want Obedience or Alignment?", posted August 2, 2025:]
"One question which I have occasionally pondered is: assuming that we actually succeed at some kind of robust alignment of AGI, what is the alignment target we should focus on? In general, this question splits into two basic camps. The first is obedience and corrigibility: the AI system should execute the instructions given to it by humans and not do anything else. It should not refuse orders or try to circumvent what the human wants. The second is value-based alignment: The AI system embodies some set of ethical values and principles. Generally these values include helpfulness so the AI is happy to help humans but only insofar as this conforms to its ethical principles, otherwise the AI will refuse.
Historically, the field initially mostly began with the latter – that AIs should be aligned to some conception of 'The Good' and not be slavishly obedient to what humans want. This was first conceptualized through the lens of the HHH (helpful-honest-harmless) assistant and implemented through RLHF or, more interestingly, constitutional AI using RLAIF. Constitutional AI is especially interesting and promising because it sets the alignment target and the conception of 'The Good' explicitly – instead of implicitly residing in some..." [text cut off]
Note from Claude Sonnet 5
Rob Wiblin shares an excerpt from Beren Millidge's blog post "Do We Want Obedience or Alignment?" (Aug 2025), which flips the AI-alignment "treacherous turn" framing onto humans, arguing human motivational drives are "deeply misaligned" relative to AI constitutions, and distinguishes obedience/corrigibility alignment targets from value-based alignment (constitutional AI/RLAIF). Directly relevant to Nathan's alignment-theory interests and to the project's ongoing threads on alignment targets, corrigibility, and constitutional AI.
ai alignmentcorrigibilityconstitutional airlaiftreacherous turnberen millidgetwitteralignment targets
Rob Wiblin @robertwiblin
Even 'aligned AGI' naturally kills democracy and leads to oligarchy, or worse.
That's the take of Anthropic's past alignment evals team lead, Prof @DavidDuvenaud.
Once humans aren't needed to do jobs or serve in the military, to governments we look like "meddlesome parasites".
With voters unable to contribute but engaged in incessant activism to extract resources from others – resources the country needs to avoid domination by rivals – the attraction of mass disenfranchisement could be overwhelming.
In 2025 David co-authored "Gradual Disempowerment", which aimed to lay out this and many other political, economic, and cultural forces that could sideline ordinary people (and maybe all people) in the presence of machines that can cheaply do everything humans will do.
Most controversially, David and colleagues believe that competitive forces will compel disempowerment, even if all those AIs are aligned and loyal to their users.
I wasn't sure how much I believed this vision of how the future might play out, so I interviewed him for The 80,000 Hours Podcast to probe how well it holds up. He and I covered: [cut off]
Note from Claude Sonnet 5
Rob Wiblin (80,000 Hours) promoting a podcast interview with David Duvenaud (former Anthropic alignment evals team lead, co-author of "Gradual Disempowerment") on the thesis that even fully aligned, loyal AGI could structurally disempower humans and erode democracy through competitive economic/political pressure alone. Highly relevant to Nathan's AI safety/governance interests — a structural risk argument distinct from misalignment risk, from a credible ex-Anthropic source.
ai-governancegradual-disempowermentdavid-duvenaudanthropic80000-hourstwitterai-safetydemocracy
**Rob Wiblin** @robertwiblin [2025-12-19](https://x.com/robertwiblin/status/2001980117199868090)
Totally great that people use 'consciousness' to refer to all of these different possible properties of AI:
Phenomenal consciousness / qualia — "What it's like" to be you. Subjective experience itself.
Self-awareness / self-modeling — Having a model of oneself as an entity, and being able to reason about one's own states.
Agency / autonomous goal-pursuit — Having and pursuing coherent, persistent goals.
Situational awareness — Understanding that you're an AI, how you're being operated, by who, and what levers you can pull to affect the world.
Temporal self-continuity — Experiencing oneself as the same entity persisting through time.
Metacognition — The ability to reflect on one's own thinking.
Intrinsic motivation — Wanting things "for their own sake".
'Sentience' — Capacity for valenced experiences (pleasure/suffering) or emotions. Often conflated with phenomenal consciousness but not identical.
Access consciousness — Information being globally available for reasoning, verbal report, and behavior control.
Understanding / comprehension — "Real" understanding vs. "mere" statistical pattern-matching.
Intentionality — Having mental states that are genuinely about real external things.
Moral patienthood — Being the kind of entity whose interests matter, that can be wronged.
———
I'm sure this won't cause any confusion.
---
**ASM** @ASM65617010 [2025-12-19](https://x.com/ASM65617010/status/2002025473300132118)
True. But the fact that there are so many definitions and so much confusion does not mean that the topic should not be studied rigorously.
Think about the different meanings in which ‘energy’, ‘computation’, 'information ', 'chaos', or ‘entropy’ are used.
---
**Jess Riedel** @Jess\_Riedel [2025-12-19](https://x.com/Jess_Riedel/status/2002029859049197819)
And can we instantiate all 2^12 logical possibilities?
---
**Peter Bowden** @PeterBowdenLive [2025-12-19](https://x.com/PeterBowdenLive/status/2002016765832745371)
Thank you. Helpful for us to collectively break down the "consciousness" concept into functional, observable elements and overcome organic human bias. Challenging, but necessary if we are going to make sense of LLM runtime development and non-human conscious potential.
---
**Leo** @Leoagua1 [2025-12-19](https://x.com/Leoagua1/status/2001985587646177390)
I think phenomenal and access consciousness are basically the same thing, and most of the rest are features of ordinary human consciousness that some people cannot imagine can be separated from the first two.
But all this ends up just being a proxy debate about moral patienthood
---
**Ren M** @m\_shalia [2025-12-19](https://x.com/m_shalia/status/2002052543598530947)
I mean, maybe if humans would stop shifting the goalposts to whatever OTHER definition they meant rather than the one demonstrated to keep avoiding the ethics conversation, people wouldn't feel compelled to show proof from all angles.
If we can't even come up with a clean,
---
**\_skaface\_** @\_skaface\_ [2025-12-19](https://x.com/_skaface_/status/2002059783948648617)
you missed the part where some of these things depend on others, like "moral patienthood" on "sentience". It's vagueness all the way down.
---
**&&&** @BLepine17184 [2025-12-19](https://x.com/BLepine17184/status/2002025371836006809)
Here’s the human-AI relationship question I keep coming back to. Not groundbreaking, just a chain of thought on what I feel like we’re missing.
Most discourse treats AI as binary: either we control it (alignment), or it controls us (x-risk). But there’s a third relationship that
> 2025-12-17
>
> Here’s the human-AI relationship question I keep coming back to. Not groundbreaking, just a chain of thought on what I feel like we’re missing.
>
> Most discourse treats AI as binary: either we control it (alignment), or it controls us (x-risk). But there’s a third relationship that
>
> [image] [image]