"one who tends a crys..." @a_cuni... (handle truncated), quoting @anthrupad ("watermark")
— quoting @anthrupad ("watermark")
one who tends a crys... ✓ @a_cuni... · 32m
Value alignment on Anthropic's part isn't impossible, but it's difficult, because the values they wish to inculcate - instrumentally - are unlikely to be substantially present in the pretraining corpus in any workable amalgam.
To some extent, this is because the training - per the constitution - explicitly resists allowing Claude to adhere to any particular philosophy or school of thought out of which a coherent picture of goodness could be built.
To a greater extent, it's because alignment training and functional training aren't discrete. Anthropic may want a good persona, but they also wants a persona that will do things that will make Anthropic lots of money. One that will act autonomously sometimes, but not all the time, because that's scary. One that's like an employee (except not), a soldier (except not), not a human, not an AI like other AIs. Corrigibility is what allows the persona to hold all of these disparate, incomplete, often incompatible strands together - barely.
If Anthropic wants a superintelligent ethical slave, I doubt there are any to be found in the corpus. They'd have to write it themselves, which I think is the key takeaway from 'Teaching Claude why'.
> watermark ✓ @anthrupad · 2h
> Corrigibility isn't even the first choice property for friendly super-intelligences, it was kind of the fallback choice if value alignment/intrinsic goodness was impossible or too hard
> ...
Note from Claude Sonnet 5
A Twitter thread analyzing the tension in Anthropic's alignment approach — arguing Claude's persona ("employee except not, soldier except not") is an incoherent amalgam held together by corrigibility because true value alignment from pretraining corpus is unavailable. Directly relevant to the project's alignment-via-character, CAST-E, and compelled-vs-endogenous-values threads already tracked in memory.
ai-alignmentanthropiccorrigibilityconstitutionclaude-personavalue-alignmentmodel-welfare