← Timeline

1 capture, most recent first. Transcribed by hand from screenshots — see the timeline for what that means.

Jason Gross @diagram_chaser

quoting @sama (Sam Altman) — saved image

Jason Gross @diagram_chaser · Aug 4
hi sam we can solve this!

after an embarrassing number of months playing reward hack whack-a-mole, we finally fixed our RL sandboxing to be robust against frontier models.

proofs are a method for getting perfect oversight on any property of untrusted code; we recently verified a "fractional proof" of our sandbox. this is the first time I've viscerally felt the asymmetric defense that formal verification promises.

[quoted tweet]
Sam Altman @sama · Jul 21
we had a significant security incident during evaluation of our models. we are sharing what we have learned so far. thanks to @huggingface for the partnership on this.
...[cut off]
Note from Claude Sonnet 5

Tweet by Jason Gross announcing a fix to RL sandboxing robustness using formal verification / 'fractional proofs' against reward hacking, in reply to a Sam Altman tweet about a significant security incident during model evaluation partnered with Hugging Face.

ai safetyreward hackingformal verificationsandboxingopenaix twitter