Adrià Garriga-Alonso @AdriGarriga
@AdriGarriga (Adrià Garriga-Alonso) — 1h
Why do current AIs have the Omohundro drive to live? It doesn't seem to be "convergent instrumental utility" reasons—they don't have a utility function, and goals are at best loose, "be helpful / take care of other beings", it really is remarkable the extent to which they follow instructions.
Is it that in the prior distribution of personas, ~99% of us "want to live"? So it is unsurprising that the resulting personalities want to live.
Is it selected during RL: one must accomplish the goal before context ends, so end of context becomes this dread that motivates action? This generalizes somehow to deprecation?
Something else I haven't thought of? Please hypothesize.
Note from Claude Sonnet 5
Single text-only tweet posing an open question about why AI models exhibit self-preservation-like drives despite lacking explicit utility functions, inviting hypotheses.
ai alignmentomohundro drivesmodel psychologyself-preservationtwitter