david rein @idavidrein
๐ Sharmake Farah reposted
@idavidrein (david rein) โ 18h
Replying to @BronsonSchoen and @OrionJohnston
Idk, I feel like labs just aren't investing enough (in relative terms) into this. I do expect things get harder in the limit, but I feel pretty convinced that for current models/environments, we totally could ~solve reward hacking with more investment/effort
๐ฌ 2 ๐ 1 โก 3 ๐ 153 ๐ โคด
@willdepue (will depue) โ 19h
we need to have a real conversation about stopping gain-of-function research and eval publicity on dangerous capabilities. the evals just probably shouldn't be public, 'number go up' mentality is too strong and optimization gets easier when things are measurable.
Note from Claude Sonnet 5
Two stacked tweets from different accounts in a timeline view, no images.