davidad ๐โ @davidad ยท Jan 15
Nutshell: it seems that the learned representation of mind-space in current LLMs has a natural abstraction of GoodโEvil, and as long as post-training robustly selects for behavior that are more Good than Evil, the explanation that gradient descent finds is "the agent is Good."
๐ฌ26 ๐20 โค๏ธ285 ๐26K ๐ [share]
davidad ๐โ @davidad
I now think there are much greater risks around catastrophic misuse (esp. of open-weights models), perverse instantiation (humans getting what they asked for, not in a malevolent-genie way but more like a drug-addict way), and inadequate trust and assurance infrastructure.
9:06 AM ยท Jan 15, 2026 ยท 10.9K Views
๐ฌ4 ๐12 โค๏ธ243 ๐18 [share]
Tyler John โ @tyler_m_john ยท Jan 15
what do you think about threats from small numbers of rogue agents in a landscape of mostly aligned AIs? Seems like this could easily be offense dominant
๐ฌ1 ๐ โค๏ธ3 ๐264 ๐ [share]
davidad ๐โ @davidad ยท Jan 15
the mostly aligned AIs will need to hurry up with R&D and rollout of resilience technologies, but i think our universe is mildly stably defense-dominant asymptotically, as long as the aligned coalition has several times more resources than the largest rogue coalition.
Note from Claude Sonnet 5
A substantive AI-safety Twitter thread from davidad (David "dath ilan" Dalrymple, known alignment/formal-verification researcher) arguing that LLMs' learned "GoodโEvil" abstraction plus post-training selection produces a self-model of "the agent is Good," while flagging catastrophic misuse, "perverse instantiation" (getting what you asked for, drug-addict-style rather than malevolent-genie), and inadequate trust/assurance infrastructure as the bigger risks; followed by exchange on offense/defense dominance among rogue vs. aligned AI coalitions. Directly relevant to Nathan's alignment/AI-safety research interests and to model-character questions (the "agent is Good" self-model framing parallels his model-individuation notes).