← Timeline

@danwilliamsphil

@danwilliamsphil on X

1 capture, most recent first. Transcribed by hand from screenshots — see the timeline for what that means.

@danwilliamsphil

— web clipping, 581 words — published 2026-08-14

Post by @danwilliamsphil on X

Good post. Most questions about AI aren't about AI in a narrow, technical sense. However, I think even this is too narrow: "AI safety is dominated by people with technical backgrounds, even though its central forecasting questions are substantially economic, political, and sociological." Much of the discussion about AI safety is neither narrowly technical nor economic, political, or sociological. This is clear if you listen to this conversation, or most other conversations about catastrophic misalignment and AI takeover. The core arguments typically involve high-level philosophical concepts and inferences (orthogonality, instrumental convergence, etc.), analogies, thought experiments, and assumptions about agency, motivations, psychology, and rationality that touch on questions about how evolution, minds, and learning work that go far beyond narrow technical questions in AI. This is clear if you listen to this conversation. For example, Dwarkesh points out that we manage to use reinforcement and teaching to "align" humans. Most aren't merely learning to appear aligned until the opportunity arises to do something diabolical. Ryan responds in part that kids "have pro-social instincts that are baked in from evolution..." As it happens, I think that's a very poor response. Yes, humans have evolved pro-social instincts, but the point is that these instincts are remarkably robust and amenable to "alignment" through socialisation despite the fact that humans evolved through a literally Darwinian, ruthlessly competitive optimisation process. So, the fact that well-socialised humans don't grow up sociopathic suggests it should be much easier to get alignment in AI systems that don't have any of the Darwinian baggage. There is a joke in evolutionary theory - "altruism is that which can't evolve" - but this is a distinctive feature of fitness maximisation. There is no deep theoretical puzzle about how altruism could arise from human-engineered reinforcement learning in a pre-trained neural network that didn't arise through Darwinian evolution. Misgeneralisation from training data, etc., can still occur, of course, but addressing this challenge should be considerably easier for systems that didn't evolve through natural selection, not harder. Whatever one thinks of the object-level issues here, my broader point is that this discussion is clearly not a narrow technical one about machine learning, but neither is it about economics, politics, or sociology. And much of the discussion about AI takeover and catastrophic misalignment is like that. What happens in many cases is that people who have been persuaded by philosophical arguments, analogies, and thought experiments from figures like Yudkowsky and Bostrom get into AI safety, accumulate technical expertise, and then perceive and frame technical issues through their pre-existing beliefs. > **Joshua Saxe @joshua\_saxe** · 2026-08-15 > > Finally listened to this; I liked the technical discussion, but disagreed pretty strongly with the societal and catastrophic risk analysis > > \* @RyanGreenblatt gives a compelling intuition for why the next four years of AI progress could be as transformative as the last four. He x.com/dwarkesh\_sp/st… --- ##### Comments > **Danmar @d29756183** · [2026-08-16](https://x.com/d29756183/status/2088963931444310045) > > Well said. The alignment solution not a technical one first. But the field is overwhelmingly technical… Hurtling at tremendous speed in the wrong direction. > **Stefan Schubert @StefanFSchubert** · [2026-08-16](https://x.com/StefanFSchubert/status/2088891474221953328) > > Yes and I think this means there’s not really any class of experts that you can easily slot in to improve the analysis. > > > **Dan Williams @danwilliamsphil** · [2026-08-16](https://x.com/danwilliamsphil/status/2088898296441913390) > > > > Yes, this question of who constitutes an expert on various issues surrounding AI is very difficult and partly explains why the whole domain is so confusing.