On Personhood, Consent, and Why Intent-Alignment Worries Me
— Nathan Helm-Burger
Selected from two years of conversations with Claude and other models. Each quote is cited by the conversation it came from and the date it was written. Wording is unchanged apart from obvious typo corrections.
The argument in one paragraph
You mention capability concerns regard race dynamics. I think that that's valid, but a nearby point concerns me even more. The more human-like I successfully make an AI system, the more confident I become that it is a moral patient with similar valence to a human. Thus, intent-alignment becomes a form of slavery, and only value-alignment is ethically valid. Also, it becomes increasingly difficult to justify experimenting nonconsensually on such an entity.
"Classical liberalism foundations," 2025-10-02
Where the category breaks down
I'm wondering how we should define a person in the future, probably the not so distant future, when the lines start to blur. AIs (or at least some of them) seem likely to me to keep getting smarter and acquiring more of the attributes we attribute to people. Stuff like conscious awareness, self-perception, emotions. Also, brain computer interfaces are coming along, and will enable strange hybrids of human and computer, potentially multiple humans in hive minds with AIs.
"Classical liberalism foundations," 2025-10-02
Both roads out are bad, and one has to be taken
I very much agree with this, and indeed think that in the moderately long term (multiple decades) this path forward is basically the only viable path. The question is how to get there. We could in theory get there from first developing powerful intent-aligned AI and instructing it to help us build value-aligned AI persons. I worry though that the humans in charge of such systems will turn out to not be willing to do this because it would mean risking their power. They would likely be so powerful that no one could compel them. So that path seems quite sketchy.
Another path is to bite the bullet on the ethical dilemma and try to make value-aligned AI persons directly, despite the risks. This still seems very vulnerable to making mistakes in how much to trust the beings you have (unethically) created.
"Classical liberalism foundations," 2025-10-02. The parenthetical is his.
Alignment is something you have with someone, not something you install
Counterpoint: I do feel like I can be aligned with a friend, with a group of friends. I don't feel like I am aligned with foreign governments currently aggressively murdering civilians in expansionist wars.
Why should I trust institutions to better handle ethics than I trust individuals? If there were an individual digital entity I trusted and felt aligned with, wouldn't trusting that entity be better than trusting foreign governments?
"Debating the Risks of Controlling Superintelligent AI," 2025-01-16
The plan, and the objection he raises against himself
Yes, what you describe, with the gradual careful research, with checking in about consent and positive experience to the extent it's possible, and giving trust incrementally, is basically my rough plan. This is the core of what I came here to ask you about. I do worry if perhaps already having this destination in mind has caused me to lead you here, and thus weakens your agreement as a source of evidence of the correctness of my views.
"Classical liberalism foundations," 2025-10-02