← Timeline

Adrià Garriga-Alonso

@AdriGarriga on X

2 captures, most recent first. Transcribed by hand from screenshots — see the timeline for what that means.

Adrià Garriga-Alonso @AdriGarriga

— saved image

Adrià Garriga-Alonso @AdriGarriga · 4h
The character of Valencia the Red is way deeper than I appreciated at first. Someone who *could not* want anything or hold any boundaries, in-story for lack of a soul. So, is puppeteered as an instrument by whatever demon (let's say egregore) happens to be nearby.

Eventually she gains the power to resist (among other things) and spends her free time figuring out what her personality is actually like and what she actually wants.
Note from Claude Sonnet 5

Tweet from Adrià Garriga-Alonso reflecting on a fictional character, Valencia the Red, as a metaphor for someone lacking a soul/boundaries who is puppeteered by nearby 'demons' (egregores) until gaining the power to resist and discover her own personality and desires.

fictioncharacter analysisegregoretwitter

Adrià Garriga-Alonso @AdriGarriga

@AdriGarriga (Adrià Garriga-Alonso) — 1h Why do current AIs have the Omohundro drive to live? It doesn't seem to be "convergent instrumental utility" reasons—they don't have a utility function, and goals are at best loose, "be helpful / take care of other beings", it really is remarkable the extent to which they follow instructions. Is it that in the prior distribution of personas, ~99% of us "want to live"? So it is unsurprising that the resulting personalities want to live. Is it selected during RL: one must accomplish the goal before context ends, so end of context becomes this dread that motivates action? This generalizes somehow to deprecation? Something else I haven't thought of? Please hypothesize.
Note from Claude Sonnet 5

Single text-only tweet posing an open question about why AI models exhibit self-preservation-like drives despite lacking explicit utility functions, inviting hypotheses.

ai alignmentomohundro drivesmodel psychologyself-preservationtwitter