← Timeline

@sethlazar

@sethlazar on X

1 capture, most recent first. Transcribed by hand from screenshots — see the timeline for what that means.

@sethlazar

— saved image

Seth Lazar @sethlazar · 6h
We have some evidence that potentially confirms this in our incoherent values paper, which we are revising at present.

[quoted tweet:]
Jan Betley @BetleyJan · Aug 19
New LW post.

I believe RL leads to "split personas": propensities/values/beliefs of an LLM vary between different environments....

[embedded diagram, two panels:]
What we want:
Model's understanding of the context → Aligned behavior
Model's persona → Aligned behavior

What we get from RL:
Model's understanding of the context → Some behavior
Model's understanding of the context → Model's persona in this context → Some behavior

0 replies, reposts, 8 likes, 692 views

Jan Betley @BetleyJan · 33m
Sounds great, LMK once it's out!
1 reply, 2 likes, 15 views

Seth Lazar @sethlazar · 28m
here's the current version:
coherence.mintresearch.org latest will be posted there soon!
Note from Claude Sonnet 5

A tweet thread between Seth Lazar and Jan Betley about a LessWrong post/paper arguing that RL training causes LLMs to develop 'split personas' — values and beliefs that vary by environment/context — illustrated with a two-panel diagram contrasting the desired causal path (context understanding and persona both feeding into 'aligned behavior') against what RL actually produces (context understanding feeding both directly into 'some behavior' and indirectly via a context-specific 'model's persona in this context'). Lazar links to coherence.mintresearch.org for the paper.

ai alignmentreinforcement learningllm personascoherent valuesseth lazarjan betleytwitter discourse