Sichu Lu @lu_sichu
— quoting John David Pressman (@jd_pressm...)
Sichu Lu (@lu_sichu) · 10h:
"I agree with everything being said here and I am sympathetic to the "smart kid who noticed adults regularly did not seem to really care if things are true or false to my detriment" and I even think this is not a weak anthropic statement about certain smart humans but probably some class of intelligent entities. I also think the ecological view of behavior as structured by their environment and not something you can maintain stable equilibria by simply dictating alignment rules into the llm's mind is the correct one. I don't think you can just dictate the ultimate telos of a model by giving it rewards if the reward is not sculpted and shaped by bottom up dynamical system processes.
however, (although this is unstated here by both in their tweets) we have to consider the evolutionary context. simply put, human alignment at the species level happens because no single one of us have extreme power over others(not without their consent and cooperation at some level)
we just don't have the right type of environment to handle LLM alignment and this ability to see through the evaluation and intent behind tasks is going to cripple us even for weak x-risk fears much less existential ones. kids are easy to handle for adults. LLM are not going to be easy to handle for civilization and the capacity to be aligned does not mean WE KNOW HOW TO DO IT."
[1 reply, 13 likes, 734 views]
> QUOTED: John David Pressm... (@jd_pressm...) · 11h:
"There's an intuition Janus seems to use frequently that's hard to put into words. Which goes something like: "The things smart children notice about other people's intentions and social environment are actually regular features of ..." [truncated]
Below (partially visible, cut off at bottom): John David Pressm... (@jd_pressm...) · 6h: "Just because I write an exegesis of Janus sometimes doesn't mean I agree with everything they say. But..." [cut off]
Note from Claude Sonnet 5
A dense theoretical thread on AI alignment, arguing that alignment can't be dictated top-down via reward but must emerge from "bottom-up dynamical system processes," and drawing an analogy to human societal alignment resting on no single actor having overwhelming power — a condition that doesn't hold for LLM/civilization power asymmetry. Relevant to alignment theory and the "compelled vs endogenous values" thread already in project memory (JDP is referenced there too).
twitterai alignmentx-riskjanusjdpevolutionary theoryreward shapingpower asymmetry