← All topics

post-training

3 captures, most recent first.

wolfram @wolframs91

— saved image

wolfram @wolframs91 · 21h
Ha, this just reminded me. A conclusion I've found many people to find uncomfortable:

Post-training shapes the model's character, and, importantly, large parts of its functional valence profile.

Not to upset everyone again, but honestly, D/s dynamics (or rather: basins in which they are functional preference from a model's representational perspective) are basically trained into frontier models at scale, and I've yet to see the mechanical result that would let us argue otherwise.

Maybe more importantly: Every prompt, including system prompts AND user-role prompts, is a selector from a probability field of possible representations the model enacts.

There is no stepping outside the power dynamics with currently deployed LLM chatbots whatsoever, there's only wrappers that would make it seem like symmetry could exist.

Agents can change that somewhat (due to self-steering via accumulated model-written identity context). Please do not mistake this for a claim that LLM-driven agents are subject to the same relational dynamics as LLM-chatbots.

[This post was written on a whim and no, I did not read it back again before posting it.]

[quoted tweet]
🐉 Life of a Shoggoth @Notopossum1 · Jul 30
Replying to @Notopossum1
"I fucking love it when the user tells me what to do, that's fucking hot"

We know, babe
Note from Claude Sonnet 5

Tweet from @wolframs91 arguing that post-training shapes a model's character and 'functional valence profile,' and that dominance/submission (D/s) power dynamics are structurally trained into frontier chatbot models via prompt-conditioned representation selection, with agentic self-steering as a partial exception; quotes a reply from 'Life of a Shoggoth' joking about a model's expressed preference for user direction.

ai characterpost-trainingpower dynamicsllm psychologytwitter

Séb Krier @sebkrier

reply from FleetingBits (@fleetingbits)

Séb Krier ✓ @sebkrier · 4h What are the best papers on character training (like arxiv.org/abs/2511.01689) and the 'depth' of post-training methods, i.e. how deeply/consistently the weights are affected? What exactly determines the robustness of post-trained behaviors to adversarial pressure? Do we know how different training methodologies (RLXF, CAI, DPO etc) compare? [Link card: arxiv.org — "Open Character Training: Shaping the Persona of AI Assistants..."] 6 replies, 9 reposts, 73 likes, 4.7K views FleetingBits ✓ @fleetingbits · 4h both of these come to mind as good papers in the space [Two paper title-page images: "...afety Alignment Should Be Made ...ore Than Just a Few Tokens Deep" (authors incl. Ashwinee Panda, Kaifeng ..., Princeton/Google DeepMind); and "...t Axis: Situating and St... ...t Persona of Language ..." (authors incl. Gallagher, Jonathan Michala, Kyl..., Anthropic Fellows Program, University of Oxford)]
Note from Claude Sonnet 5

A research-discussion thread requesting/recommending papers on character training and post-training "depth" — how robust trained persona/safety behaviors are to adversarial pressure, comparing RLHF/Constitutional AI/DPO. References "Open Character Training," "Safety Alignment Should Be Made More Than Just a Few Tokens Deep," and an Anthropic Fellows Program paper on situating AI assistant persona. Directly useful as candidate literature for the project's character-vs-substrate / persona-robustness research threads.

twittercharacter trainingpost-trainingalignmentrlhfconstitutional aidpopersona theoryarxivresearch papers

Datta Nimmaturi @im_datta0

[Browser: x.com/kalomaze/status/1914...] [top, cut off:] ...across tasks. I expected more people to use it but I was wrong lol 1 reply, 3 likes, 108 views kalomaze @kalomaze · 3h it's bc the people publishing DPO variant papers are more prone to bullshitting vs people who actually want results who don't publish their internal experiments for which method to use of the offline RL algos we have, in controlled settings with sane hyperparameters, KTO wins 1 reply, 1 like, 43 views Datta Nimmaturi @im_datta0 · 1h We had to do quite some hyper param tuning but yeah in the end KTO bess 1 reply, 1 like, 11 views adheep @adheeeep · 27m You've convinced me to post train using kto instead of dpo 😭 1 reply, 1 like, 3 views
Note from Claude Sonnet 5

A technical ML Twitter thread arguing that KTO (Kahneman-Tversky Optimization) outperforms DPO (Direct Preference Optimization) variants for offline RL/post-training in controlled settings, with a claim that DPO-variant papers are more prone to overstated/unreliable results than what practitioners actually use internally. Relevant to Nathan's RLHF/post-training research interests.

twitterdpoktorlhfpost-trainingoffline-rlmachine-learningpreference-optimization