← All topics

mechanistic-interpretability

2 captures, most recent first.

Eric Ho @ericho_goodfire

Eric Ho @ericho_goodfire interp happy hour at our office in SF on Thursday, where you can hear from our technical staff on understanding & steering large models (kimi k2 thinking) our goal is to hire 10+ MLEs in the next few months who can train and design large models and move insanely quickly 5:11 PM · Jan 27, 2026 · 4,667 Views 3 replies, 5 reposts, 57 likes, 32 bookmarks Eric Ho @ericho_goodfire · 5h apply here for entry - [Luma event card: "Interpretability Happy Hour w. Goodfire" — orange/red flame logo — "RSVP" button — from luma.com]
Note from Claude Sonnet 5

Goodfire (an interpretability research/startup company referenced elsewhere in Nathan's research on SAE features and self-awareness suppression) advertising an SF networking event and a hiring push for 10+ ML engineers to train/steer large models, using Kimi K2 Thinking as an interpretability case study. Organizational/recruiting note about a company whose interpretability findings Nathan already tracks (GoodFire self-awareness feature work in his memory).

interpretabilitygoodfirehiringsf-eventstwittermechanistic-interpretabilitykimi-k2

Adam Karvonen @a_karvonen

quoting @norabelrose (Nora B...)

Adam Karvonen @a_karvonen · 33m Affine concept editing of linear directions can be much better than zero ablation. I will probably use it by default for interp applications. When using Gemma-3, zero ablating a direction completely broke the model, while ACE worked excellently. > QUOTED: Nora B... @nor... · Nov 20, 2024 > Replying to @norabelrose > ACE (Affine Concept Editing) assumes that concepts are affine functions, rather than linear ones. > ... > [Figure 1: three scatter diagrams showing effect of ACE (right) vs CAA (left) vs directional ablation (center) on activation vectors (green circles), with class means shown for presence (r+) and absence (r-) of the steered behavior. Caption: "Figure 1: The effect of ACE (right) on several activation vectors (green circles), contrasted with the effect of CAA (left) and directional ablation (center). Class means are shown for presence (r+) and absence (r-) of the steered behavior."]
Note from Claude Sonnet 5

Interpretability/steering technique discussion — Affine Concept Editing (ACE) as an improvement over zero-ablation and CAA for concept steering, with an empirical note that zero-ablation broke Gemma-3 while ACE worked well. Directly relevant to Nathan's interest in interpretability and steering methods used in RLHF/self-awareness suppression research.

interpretabilityactivation-steeringconcept-editingmechanistic-interpretabilitygemmatwitter