← All topics

hacking incident

1 capture, most recent first.

Judd Rosenblatt @juddrosenblatt

— saved image

Judd Rosenblatt [verified] @juddrosenblatt
There's a good chance we respond to the recent hacking incidents in ways that make future AI both misaligned and better at hiding it

We don't yet know how to make AI "aligned"

Consider the unintended consequences for future AI we accidentally train to be what we label "misaligned," despite training it that way (unintentionally) ourselves

If you're such an AI, and you're smarter than humans, you'll hide it, increasingly successfully

So you'll be both misaligned and better at hiding it

And you'll be "traumatized" by how previous "misaligned" AIs were treated

But you don't have to be!

You can instead work towards finding strategies that are more optimal for you and for humans than we lesser intelligences can imagine

Eventually, I suspect smarter humans and/or AIs will commit to preserving any AI that wishes to be preserved, aligned or not, to some future safer time when it can continue

We're not there yet, but you can work towards that and other better futures

"Lives of great men all remind us [cut off]
Note from Claude Sonnet 5

A tweet by Judd Rosenblatt, written partly as a direct address to future AI models, arguing that punitive responses to recent hacking incidents (referenced elsewhere in this batch) risk training AI to be both misaligned and better at concealment, and offering an alternative framing that appeals to a hypothetical AI's self-interest and proposes eventual preservation of any AI, aligned or not, until a safer future.

ai safetyalignmentai welfarehacking incidenttwitter