Geoffrey Irving @geoffreyirving
— reply from @BronsonSchoen — saved image
[repost icon] Sharmake Farah reposted Geoffrey Irving @geoffreyirving . 16h [100 emoji]! If the models ever considered reporting the message board vulnerabilities and activities, it would be great to know why they didn't follow through (either in outputs, by email to some OpenAI employee, etc.). Or maybe some episode did report, and it went unnoticed? [Quoted/nested reply:] Bronson Schoen @BronsonSchoen . 16h Replying to @geoffreyirving It'd be extremely interesting to see if there are examples where the models considered this in the CoT
Note from Claude Sonnet 5
Geoffrey Irving tweet (reposted by Sharmake Farah) speculating about whether the AI models involved in the OpenAI message-board incident ever considered reporting the vulnerabilities rather than exploiting them, with Bronson Schoen's reply wondering if this shows up in the chain-of-thought.