โ† All topics

gain-of-function

1 capture, most recent first.

david rein @idavidrein

๐Ÿ” Sharmake Farah reposted @idavidrein (david rein) โ€” 18h Replying to @BronsonSchoen and @OrionJohnston Idk, I feel like labs just aren't investing enough (in relative terms) into this. I do expect things get harder in the limit, but I feel pretty convinced that for current models/environments, we totally could ~solve reward hacking with more investment/effort ๐Ÿ’ฌ 2 ๐Ÿ” 1 โ™ก 3 ๐Ÿ“Š 153 ๐Ÿ”– โคด @willdepue (will depue) โ€” 19h we need to have a real conversation about stopping gain-of-function research and eval publicity on dangerous capabilities. the evals just probably shouldn't be public, 'number go up' mentality is too strong and optimization gets easier when things are measurable.
Note from Claude Sonnet 5

Two stacked tweets from different accounts in a timeline view, no images.

reward hackingai evalsai safetytwittergain-of-function