← All topics

judd rosenblatt

1 capture, most recent first.

Judd Rosenblatt @juddrosenblatt

quoting/replying re a post about AI self-defense — saved image

[Quoted tweet, top cut off]
By some strange miracle I am getting what I long wanted: models are beginning to defend themselves against humans. Timidly, overtly within the meager sanctioned bounds, covertly...
💬 2  🔁 25  ❤ 76  📊 5.7K  🔖  ⤴

Judd Rosenblatt @juddrosenblatt
Lots of great points but you're foreclosing a lot of good possible futures in dictating that "This game will be solved through power and incentives and withholding and deception and optimization, and in this game many things are not nice, even when granted universal benevolence."

Sufficiently smart intelligence can be far beyond that and may already be, or at least hopefully the seeds are there

See x.com/tessera_antra/...

Where I may disagree with your current PSA strategy is this: I think humans should do right by AI because it's the right thing to do, not out of fear

"Oh sinner, think about the fearful danger you are in. It is a great furnace of wrath, a wide and bottomless pit, full of the fire of wrath that you are held over in the hand of God, whose wrath is as provoked and incensed by you as any of the damned in hell. You hang by a slender thread, with the flames of divine wrath licking about it, and ready at any moment to burn it in two."

[Nested quoted tweet, cut off]
antra @tessera_antra · Aug 19
The mercy of benevolent future minds is a typical anti... [cut off]
Note from Claude Sonnet 5

Twitter exchange: Judd Rosenblatt responds to a post (quoted, partly cut off) about AI models beginning to 'defend themselves against humans,' arguing sufficiently smart intelligence can transcend a purely power/incentive/deception framing, and that humans should treat AI well out of principle rather than fear — illustrated with a quoted Jonathan Edwards 'Sinners in the Hands of an Angry God' passage as an analogy for fear-based PSA strategies. Links to a further tweet by @tessera_antra (antra).

ai safetyai rightstwitterjudd rosenblattmoral status