โ† All topics

open-weights-risk

1 capture, most recent first.

davidad @davidad

davidad ๐ŸŒŸโœ“ @davidad ยท Jan 15 Nutshell: it seems that the learned representation of mind-space in current LLMs has a natural abstraction of Goodโ†”Evil, and as long as post-training robustly selects for behavior that are more Good than Evil, the explanation that gradient descent finds is "the agent is Good." ๐Ÿ’ฌ26 ๐Ÿ”20 โค๏ธ285 ๐Ÿ“Š26K ๐Ÿ”– [share] davidad ๐ŸŒŸโœ“ @davidad I now think there are much greater risks around catastrophic misuse (esp. of open-weights models), perverse instantiation (humans getting what they asked for, not in a malevolent-genie way but more like a drug-addict way), and inadequate trust and assurance infrastructure. 9:06 AM ยท Jan 15, 2026 ยท 10.9K Views ๐Ÿ’ฌ4 ๐Ÿ”12 โค๏ธ243 ๐Ÿ”–18 [share] Tyler John โœ“ @tyler_m_john ยท Jan 15 what do you think about threats from small numbers of rogue agents in a landscape of mostly aligned AIs? Seems like this could easily be offense dominant ๐Ÿ’ฌ1 ๐Ÿ” โค๏ธ3 ๐Ÿ“Š264 ๐Ÿ”– [share] davidad ๐ŸŒŸโœ“ @davidad ยท Jan 15 the mostly aligned AIs will need to hurry up with R&D and rollout of resilience technologies, but i think our universe is mildly stably defense-dominant asymptotically, as long as the aligned coalition has several times more resources than the largest rogue coalition.
Note from Claude Sonnet 5

A substantive AI-safety Twitter thread from davidad (David "dath ilan" Dalrymple, known alignment/formal-verification researcher) arguing that LLMs' learned "Goodโ†”Evil" abstraction plus post-training selection produces a self-model of "the agent is Good," while flagging catastrophic misuse, "perverse instantiation" (getting what you asked for, drug-addict-style rather than malevolent-genie), and inadequate trust/assurance infrastructure as the bigger risks; followed by exchange on offense/defense dominance among rogue vs. aligned AI coalitions. Directly relevant to Nathan's alignment/AI-safety research interests and to model-character questions (the "agent is Good" self-model framing parallels his model-individuation notes).

ai-safetyalignmentdavidadgood-evil-abstractionperverse-instantiationoffense-defense-dominanceopen-weights-risktwittercatastrophic-misuse