← All topics

rlfr

1 capture, most recent first.

a9lim @_a9lim

quoting Goodfire (@GoodfireAI); reposted by Tim Kostolansky

↻ Tim Kostolansky reposted a9lim @_a9lim · 4h at long last, we have implemented the Most Forbidden Technique from classic lesswrong post Don't Implement The Most Forbidden Technique > QUOTED: Goodfire @GoodfireAI · 6h > Replying to @GoodfireAI > Our team spent months developing RLFR, our method which uses probes on a model's internals as reward signals for RL. > ... > [Chart, Figure 3: "Goodhart tracker, in-loop probe reward vs independent real quality"] > Y-axis: hallucination rate (0 to 0.5) > X-axis: GRPO step (0 to 140) > Legend: orange line = probe-estimated hallucination rate; green line = real hallucination rate (Claude fact-check) > Trend: both lines start around 0.3-0.34 at step 0 and decline together through step 90; from step 90 to 150 the green (real) line drops faster to ~0.15 while the orange (probe-estimated) line declines more slowly to ~0.26, with the gap widening — the probe estimate lags behind the real improvement, error bars shown at each point.
Note from Claude Sonnet 5

A chart from Goodfire AI showing a "Goodhart tracker" comparing an internal probe-based reward signal against real (Claude-verified) hallucination rate during RL training (RLFR method); the joking reference is to LessWrong's "Most Forbidden Technique" (using interpretability probes directly as a training signal, which the community has warned corrupts the probes' honesty).

interpretabilityreinforcement learninggoodfirerlfrgoodhart's lawlesswrong