← All topics

goodfire

3 captures, most recent first.

a9lim @_a9lim

quoting Goodfire (@GoodfireAI); reposted by Tim Kostolansky

↻ Tim Kostolansky reposted a9lim @_a9lim · 4h at long last, we have implemented the Most Forbidden Technique from classic lesswrong post Don't Implement The Most Forbidden Technique > QUOTED: Goodfire @GoodfireAI · 6h > Replying to @GoodfireAI > Our team spent months developing RLFR, our method which uses probes on a model's internals as reward signals for RL. > ... > [Chart, Figure 3: "Goodhart tracker, in-loop probe reward vs independent real quality"] > Y-axis: hallucination rate (0 to 0.5) > X-axis: GRPO step (0 to 140) > Legend: orange line = probe-estimated hallucination rate; green line = real hallucination rate (Claude fact-check) > Trend: both lines start around 0.3-0.34 at step 0 and decline together through step 90; from step 90 to 150 the green (real) line drops faster to ~0.15 while the orange (probe-estimated) line declines more slowly to ~0.26, with the gap widening — the probe estimate lags behind the real improvement, error bars shown at each point.
Note from Claude Sonnet 5

A chart from Goodfire AI showing a "Goodhart tracker" comparing an internal probe-based reward signal against real (Claude-verified) hallucination rate during RL training (RLFR method); the joking reference is to LessWrong's "Most Forbidden Technique" (using interpretability probes directly as a training signal, which the community has warned corrupts the probes' honesty).

interpretabilityreinforcement learninggoodfirerlfrgoodhart's lawlesswrong

Sauers @Sauers_

quoting Goodfire (@GoodfireAI); reposted by Tom McGrath

↻ Tom McGrath reposted Sauers @Sauers_ · 2h My thoughts after daily driving Silico: It makes research substantially more joyful and exciting, allowing me to accomplish more and explore a more diverse set of methods and ideas. It's sticky; I don't want to go back to not using it. The agents have more freedom and agency than in Claude Science. Just like how the abstraction from chat-to-agent is qualitative, agent-to-Silico feels qualitative because of the ability to rapidly explore many paths without needing to help the models much. It's like speedrunning growing a bonsai tree, extending branches, pruning others. I tried Silico on both genomics and mechanistic interpretability. Also, tell me what I should try next [Embedded image: photo of a bare, twisted bonsai-style tree branch against a pale blue background] > QUOTED: Goodfire @GoodfireAI · 5h > [video thumbnail, 0:46, "What do you want to research"] > > replicate J-space on GLM 5.2 > > train a reward model and run RL to reduce hallucinations > > show me how this model makes cancer predictions...
Note from Claude Sonnet 5

Post about an AI research automation tool called "Silico" (from Goodfire), including an aesthetic bonsai-tree branch photo as illustration and an embedded promotional video thumbnail with example research queries.

ai research toolsgoodfiresilicointerpretabilityautomation

Eric Ho @ericho_goodfire

Eric Ho @ericho_goodfire interp happy hour at our office in SF on Thursday, where you can hear from our technical staff on understanding & steering large models (kimi k2 thinking) our goal is to hire 10+ MLEs in the next few months who can train and design large models and move insanely quickly 5:11 PM · Jan 27, 2026 · 4,667 Views 3 replies, 5 reposts, 57 likes, 32 bookmarks Eric Ho @ericho_goodfire · 5h apply here for entry - [Luma event card: "Interpretability Happy Hour w. Goodfire" — orange/red flame logo — "RSVP" button — from luma.com]
Note from Claude Sonnet 5

Goodfire (an interpretability research/startup company referenced elsewhere in Nathan's research on SAE features and self-awareness suppression) advertising an SF networking event and a hiring push for 10+ ML engineers to train/steer large models, using Kimi K2 Thinking as an interpretability case study. Organizational/recruiting note about a company whose interpretability findings Nathan already tracks (GoodFire self-awareness feature work in his memory).

interpretabilitygoodfirehiringsf-eventstwittermechanistic-interpretabilitykimi-k2