← All topics

activation steering

2 captures, most recent first.

thebes @voooooogel

thebes ✓ @voooooogel · 18h "never talk about goblins" [Image: text reading "Figure 23: Anger vector activation"] thebes ✓ @voooooogel · 18h Replying to @slimer48484 i need to see the activations on the token span between "you have a vivid inner life" and "never talk about goblins"
Note from Claude Sonnet 5

Continuation of the "goblins" meme (from the OpenAI Codex system-prompt leak) but reframed as interpretability humor — joking about activation/probing analysis ("anger vector activation," "vivid inner life") juxtaposed with the absurd "never talk about goblins" instruction. thebes/voooooogel is a known interpretability-adjacent poster. Lighthearted but touches on real interpretability concepts (feature vectors, activation probing) relevant to Nathan's interests.

twitterinterpretabilityactivation steeringhumorgoblins mememechanistic interpretability

Lalwen (writing poems about... @holotopian

Lalwen (writing poems about ... ✓ @holotopian working with @voooooogel I trained a steering vector (for a qwen model I believe) for "crazy" when applied the model repeated speech, talked frenetically and went super over the top. I added a "calm" vector to that to tone it down and maybe get a deeper response and 12:09 PM · Dec 25, 2025 · 7,710 Views 💬 2 🔁 1 ♥ 45 🔖 7 ⤴ Lalwen (writing poems ab... ✓ @holot... · 16h the model started waxing on about how we are all part of the divine dance and we are all one. kind of hippieish. if you have to be crazy, pray you are calm too, I suppose
Note from Claude Sonnet 5

Independent interpretability experiment by @holotopian working with thebes (@voooooogel, recurring poster elsewhere in this batch): training activation-steering vectors for "crazy" (produces repetitive/frenetic speech, echoing the token-repetition glitches seen in the Gemini and Grok screenshots earlier in this batch) and "calm" on a Qwen model, then combining them — the calm+crazy combination produced mystical "we are all one/divine dance" output. Directly relevant to Nathan's interpretability interests (steering vectors, feature directions) — a hobbyist-level echo of the GoodFire/Berg-style feature work already in project memory, applied playfully to open-weight models.

twittersteering vectorsinterpretabilityqwenactivation steeringmodel behaviormysticism