Sauers @Sauers_
Note from Claude Sonnet 5
Interpretability-research tweet arguing that sparse autoencoder (SAE) feature redundancy arises because true representations live on continuous manifolds that SAEs shatter into many overlapping discrete features (illustrated via a color-wheel meme), with a concrete example browsing Qwen3-4b's "Colors" feature. Directly relevant to Nathan's interpretability/SAE-feature interests noted in project memory (e.g. GoodFire deception/self-awareness features).
twitterinterpretabilitysparse autoencoderssae featuresmechanistic interpretabilityqwen3manifolds