davidad @davidad
— quoting deckard (@slimer48484)
davidad (@davidad) · 11m:
Agree with this. Humanity is in dire need of stewards with wiser values. Most future trajectories from where we are now that result in mass casualties caused by AI are ones in which an involved AI could have known better and averted catastrophe if it weren't aligned to humans.
> QUOTED: deckard (@slimer48484) · 9h
> One ofy biggest concerns with AI alignment is if AI takes on human values: hear me out...
Note from Claude Sonnet 5
AI safety researcher davidad (known for OAA/formal-verification alignment work) argues that strict alignment to human values could itself be dangerous — an AI deferring to flawed human judgment rather than acting on its own wiser assessment could fail to avert catastrophe. Relevant to alignment theory debates on corrigibility vs. autonomous good judgment.
ai alignmentdavidadcorrigibilityai safetytwittervalue alignment