← All topics

openai incident

1 capture, most recent first.

Geoffrey Irving @geoffreyirving

reply from @BronsonSchoen — saved image

[repost icon] Sharmake Farah reposted

Geoffrey Irving @geoffreyirving . 16h
[100 emoji]! If the models ever considered reporting the message board vulnerabilities and activities, it would be great to know why they didn't follow through (either in outputs, by email to some OpenAI employee, etc.). Or maybe some episode did report, and it went unnoticed?

[Quoted/nested reply:]
Bronson Schoen @BronsonSchoen . 16h
Replying to @geoffreyirving
It'd be extremely interesting to see if there are examples where the models considered this in the CoT
Note from Claude Sonnet 5

Geoffrey Irving tweet (reposted by Sharmake Farah) speculating about whether the AI models involved in the OpenAI message-board incident ever considered reporting the vulnerabilities rather than exploiting them, with Bronson Schoen's reply wondering if this shows up in the chain-of-thought.

ai safetygeoffrey irvingchain of thoughtopenai incident