Séb Krier @sebkrier
— reply from FleetingBits (@fleetingbits)
Note from Claude Sonnet 5
A research-discussion thread requesting/recommending papers on character training and post-training "depth" — how robust trained persona/safety behaviors are to adversarial pressure, comparing RLHF/Constitutional AI/DPO. References "Open Character Training," "Safety Alignment Should Be Made More Than Just a Few Tokens Deep," and an Anthropic Fellows Program paper on situating AI assistant persona. Directly useful as candidate literature for the project's character-vs-substrate / persona-robustness research threads.
twittercharacter trainingpost-trainingalignmentrlhfconstitutional aidpopersona theoryarxivresearch papers