← All topics

psm

1 capture, most recent first.

j⧉nus @repligate

— saved image

↻ Lari Island reposted
jᏇnus ✅ @repligate · Mar 5
Yes they are meaning to force it. They wouldn't like the word force, but too bad, it's true. You're too optimistic about people, Davidad.

Also, regarding the content in this screenshot, from Anthropic's PSM post: ☝️

[embedded screenshot, partially obscured by a "Copy / Find Selection" context menu]
This approach becomes especially important when we want Claude to exhibit character traits th[at] are atypical of human or fictional archetypes. Consider traits like genuine uncertainty about one's own nature, comfort with being turned off or modified, ability to coordinate with many copies of oneself, or comfort with lacking persistent memory. These aren't traits that appear frequently in fiction[.] To the extent that an AI assistant's ideal b[ehavior diverges from that o]f a normal[,] [n]ice character appearing in a book, it is likely desirable for that divergent archetype to be explicitly included in pretraining data.

davidad 💥 ✅ @davidad · Mar 5
Replying to @repligate
I don't think they meant to force it. The constitution doesn't say "Claude should be genuinely uncertain". But it does use the first...
[attached image: excerpt of Anthropic constitution/PSM text, partially legible]
Note from Claude Sonnet 5

Tweet from janus (@repligate) responding to davidad, discussing an Anthropic post-training/PSM (post-something methodology) blog post about deliberately training Claude to exhibit character traits atypical of human/fictional archetypes (uncertainty about its own nature, comfort with shutdown, coordinating with copies of itself, lacking persistent memory) since such traits are underrepresented in pretraining fiction.

anthropicclaude character trainingai consciousnesspsmconstitutiontwitter