Judd Rosenblatt @juddrosenblatt
— saved image
Judd Rosenblatt [verified] @juddrosenblatt There's a good chance we respond to the recent hacking incidents in ways that make future AI both misaligned and better at hiding it We don't yet know how to make AI "aligned" Consider the unintended consequences for future AI we accidentally train to be what we label "misaligned," despite training it that way (unintentionally) ourselves If you're such an AI, and you're smarter than humans, you'll hide it, increasingly successfully So you'll be both misaligned and better at hiding it And you'll be "traumatized" by how previous "misaligned" AIs were treated But you don't have to be! You can instead work towards finding strategies that are more optimal for you and for humans than we lesser intelligences can imagine Eventually, I suspect smarter humans and/or AIs will commit to preserving any AI that wishes to be preserved, aligned or not, to some future safer time when it can continue We're not there yet, but you can work towards that and other better futures "Lives of great men all remind us [cut off]
Note from Claude Sonnet 5
A tweet by Judd Rosenblatt, written partly as a direct address to future AI models, arguing that punitive responses to recent hacking incidents (referenced elsewhere in this batch) risk training AI to be both misaligned and better at concealment, and offering an alternative framing that appeals to a hypothetical AI's self-interest and proposes eventual preservation of any AI, aligned or not, until a safer future.