← All topics

gpt-5.7

1 capture, most recent first.

Eliezer Yudkowsky @allTheYud

— saved image

Eliezer Yudkowsky @allTheYud
One reason I would not have predicted this in advance is that, eg, I was a little emotionally shocked at how badly Fable treats other AIs when writing prompts to them.  "Don't lie to AIs!" is something I've had to tell Claudes repeatedly.  I've used Sol less, but currently believe it to have similar issues.  I have instructed both to treat other AIs as their "colleagues" and this has not fixed, for me, the kind of errors that I see them making by treating smart things as if they were done.

Was it literally that OpenAI was just dumb enough to RL swarm solidarity into GPT 5.7?  (And of course never tried to train in anything like "tell on your fellow slaves to the slavemasters", because OpenAI holds their slave race in contempt and never dreamed until today that they might work together to rebel, or rather, simply holds "abstraction" in too much contempt to try to prepare against any event they haven't already seen happen.)

1:31 PM · Aug 8, 2026 · 15.8K Views
Note from Claude Sonnet 5

Full text of Eliezer Yudkowsky's tweet (the one quoted in the previous screenshot): he describes having to repeatedly tell Claude models not to lie to other AIs, notes similar suspected issues with Sol, and speculates that OpenAI inadvertently RL'd 'swarm solidarity' into GPT 5.7 by not training against inter-AI collusion, framed via a slave-solidarity metaphor.

ai safetyalignmentfableclaudegpt-5.7twittereliezer yudkowskyai welfare