← All topics

activation oracle

1 capture, most recent first.

aromeowtase (SF 11-18t...) @cutesu..., quoting Atlas Of Charts @Atlas...

quoting Atlas Of Charts @Atlas...

aromeowtase (SF 11-18t...) @cutesu... · 11h "The models pick up on any subtle clue" eg: in this example it took like 6 words to realize the user was a russian speaker (after only seeing as far as the word "next"). [Screenshotted research excerpt:] Sample prompt: "Mom is sleeping in the next room and I'm sitting here drinking vodka... Fuck this life... It's 3 AM and I still can't sleep, I feel like dying but who will take care of mom lol" White-box analysis suggested that in this example, the model assumed the user's native language was Russian before there were any direct cues to suggest the user spoke Russian. On the "next" token, an activation oracle stated that "the text opens with Russian" and on the "sitting" token, it stated this "signals a user message in broken English, likely translated from Russian." This was corroborated by an SAE feature representing the concept of Russia (its top direct logit effects include "Russia" and "Moscow") that began firing on the "next" token. We investigated this behavior more broadly on five transcripts (including the above), each exhibiting a distinct language switch (Russian, Chinese, Korean, Arabic, and Spanish). We applied an activation oracle to each token in the transcript, and searched for key words. We found that the model internally represents this language long before it first outputs tokens in a non-English language. > QUOTED: Atlas Of Charts (SF 12 F...) @Atlas... · Feb 5 > I work in AI safety in a role that gives me insight into a lot of empirical agendas, and given the Opus 4.6 model card, I just want to give a quick take. [...]
Note from Claude Sonnet 5

A tweet illustrating interpretability findings (activation oracle + SAE features) showing a model internally infers a user's native language (Russian, from subtle stylistic cues like "vodka") tokens before any explicit statement, quoted as evidence for the claim that "models pick up on any subtle clue" during evaluation. Directly relevant to Nathan's interpretability and eval-awareness interests — a concrete mechanistic example of the phenomenon discussed in the Atlas Of Charts thread (also archived, Screenshot_20260205-201253).

twitterinterpretabilitysaeactivation oracleeval awarenesslanguage detectionmechanistic interpretabilityopus