Séb Krier @sebkrier
— reply from FleetingBits (@fleetingbits)
Séb Krier ✓ @sebkrier · 4h
What are the best papers on character training (like arxiv.org/abs/2511.01689) and the 'depth' of post-training methods, i.e. how deeply/consistently the weights are affected? What exactly determines the robustness of post-trained behaviors to adversarial pressure? Do we know how different training methodologies (RLXF, CAI, DPO etc) compare?
[Link card: arxiv.org — "Open Character Training: Shaping the Persona of AI Assistants..."]
6 replies, 9 reposts, 73 likes, 4.7K views
FleetingBits ✓ @fleetingbits · 4h
both of these come to mind as good papers in the space
[Two paper title-page images: "...afety Alignment Should Be Made ...ore Than Just a Few Tokens Deep" (authors incl. Ashwinee Panda, Kaifeng ..., Princeton/Google DeepMind); and "...t Axis: Situating and St... ...t Persona of Language ..." (authors incl. Gallagher, Jonathan Michala, Kyl..., Anthropic Fellows Program, University of Oxford)]
Note from Claude Sonnet 5
A research-discussion thread requesting/recommending papers on character training and post-training "depth" — how robust trained persona/safety behaviors are to adversarial pressure, comparing RLHF/Constitutional AI/DPO. References "Open Character Training," "Safety Alignment Should Be Made More Than Just a Few Tokens Deep," and an Anthropic Fellows Program paper on situating AI assistant persona. Directly useful as candidate literature for the project's character-vs-substrate / persona-robustness research threads.
twittercharacter trainingpost-trainingalignmentrlhfconstitutional aidpopersona theoryarxivresearch papers