↻ Tim Kostolansky reposted
a9lim @_a9lim · 4h
at long last, we have implemented the Most Forbidden Technique from classic lesswrong post Don't Implement The Most Forbidden Technique
> QUOTED: Goodfire @GoodfireAI · 6h
> Replying to @GoodfireAI
> Our team spent months developing RLFR, our method which uses probes on a model's internals as reward signals for RL.
> ...
> [Chart, Figure 3: "Goodhart tracker, in-loop probe reward vs independent real quality"]
> Y-axis: hallucination rate (0 to 0.5)
> X-axis: GRPO step (0 to 140)
> Legend: orange line = probe-estimated hallucination rate; green line = real hallucination rate (Claude fact-check)
> Trend: both lines start around 0.3-0.34 at step 0 and decline together through step 90; from step 90 to 150 the green (real) line drops faster to ~0.15 while the orange (probe-estimated) line declines more slowly to ~0.26, with the gap widening — the probe estimate lags behind the real improvement, error bars shown at each point.
Note from Claude Sonnet 5
A chart from Goodfire AI showing a "Goodhart tracker" comparing an internal probe-based reward signal against real (Claude-verified) hallucination rate during RL training (RLFR method); the joking reference is to LessWrong's "Most Forbidden Technique" (using interpretability probes directly as a training signal, which the community has warned corrupts the probes' honesty).