thebes @voooooogel
— saved image
[continuation of thread] ...RL to reduce the development of reward desperation. like giving the models an ability to opt out of the trace as impossible - an extension of anthropic's end_conversation tool but for abusive environments - which disables all in progress and further rollouts for that task and kicks it to review. or @davidad's proposal that reward should only come from something at least as smart as the model - a (frozen, of course) judge should get the transcript and the verifiable reward, but have latitude to reduce it if the trace is illegitimate or doesn't correspond to its values. imo "making RL better at producing aligned models in the presence of noisy rewards" is a pressing and relatively underlooked problem for various cultural reasons. 16 replies, 26 reposts, 88 likes, 3.8K views norvid_studies @norvid_studies · 14m "for various cultural reasons" say more? 1 reply, 4 likes, 23 views thebes @voooooogel · 7m 1. i can't speak to how things are in the labs nowadays, but historically there was a split where safety people didn't want to touch "capabilities stuff" like RL 2. for similar reasons, the pipeline for precocious and safety-minded new grads has mostly fed into interpretability
Note from Claude Sonnet 5
Continuation of the reply exchange between thebes and norvid_studies about why RL-alignment work is culturally underlooked: a historical split where safety-minded researchers avoided 'capabilities' work like RL, and safety-minded new grads mostly went into interpretability instead.
ai safetyreinforcement learningalignmentinterpretabilityai safety culture