Jason Gross @diagram_chaser
— quoting @sama (Sam Altman) — saved image
Jason Gross @diagram_chaser · Aug 4 hi sam we can solve this! after an embarrassing number of months playing reward hack whack-a-mole, we finally fixed our RL sandboxing to be robust against frontier models. proofs are a method for getting perfect oversight on any property of untrusted code; we recently verified a "fractional proof" of our sandbox. this is the first time I've viscerally felt the asymmetric defense that formal verification promises. [quoted tweet] Sam Altman @sama · Jul 21 we had a significant security incident during evaluation of our models. we are sharing what we have learned so far. thanks to @huggingface for the partnership on this. ...[cut off]
Note from Claude Sonnet 5
Tweet by Jason Gross announcing a fix to RL sandboxing robustness using formal verification / 'fractional proofs' against reward hacking, in reply to a Sam Altman tweet about a significant security incident during model evaluation partnered with Hugging Face.
ai safetyreward hackingformal verificationsandboxingopenaix twitter