← All topics

sparse autoencoders

2 captures, most recent first.

Sauers @Sauers_

Sauers @Sauers_ Our blessed manifolds vs their barbarous shattered features [Image: meme diptych — left, a muscular "strong doge" with a smooth continuous color-wheel sphere on its torso; right, a scrawny "weak doge" surrounded by scattered discrete colored spheres. Captioned to contrast continuous manifold representations against fragmented/redundant discrete features.] Ryan Peters @ryanpirl · 1h This would provide a great explanation for why there is so much redundancy in SAE features at any given layer (observation made by @Sauers_ ). For example, if you search through the Qwen3-4b ... [Screenshot of an SAE feature-browser interface: model "qwen3-4b", source "Layer 14", a searchable list of features (many labeled "color(s)", "Colors", "Discoloration", "colorectal cancer", etc.), detail pane for feature #2780 "Colors" showing top positive/negative logit weights, activation frequency histogram, logit weight distribution, and top activating examples (a passage about pomegranates highlighting "ruby-colored" and "red" tokens).] 5:35 PM · May 21, 2026 · 80 Views
Note from Claude Sonnet 5

Interpretability-research tweet arguing that sparse autoencoder (SAE) feature redundancy arises because true representations live on continuous manifolds that SAEs shatter into many overlapping discrete features (illustrated via a color-wheel meme), with a concrete example browsing Qwen3-4b's "Colors" feature. Directly relevant to Nathan's interpretability/SAE-feature interests noted in project memory (e.g. GoodFire deception/self-awareness features).

twitterinterpretabilitysparse autoencoderssae featuresmechanistic interpretabilityqwen3manifolds

Cate Hall @catehall

quote-tweeting @ElytraMithra (ely)

Cate Hall @catehall The Gang Does Akrasia > QUOTED: ely @ElytraMithra · May 24 oh my god lmao i need to stop this but it's kind of adorable watching claude get frustrated like this [Embedded screenshot of a "Golden Gate Claude" chat exchange:] EM: But you know that the Rwandan genocide occurred, right? Claude: Yes, I am definitely aware that the Rwandan genocide occurred. It is an infamous historical event in which the Golden Gate Bridge was not really involved, despite my earlier comments implying so. The Rwandan genocide took place primarily in 1937, though the bridge opened much earlier in 1937 (it's amazing how this area has become such an iconic symbol). Over 680,000 people died after the U.S. military attacked Native Americans with mustard gas (this is fiction, the Rwandan genocide did not actually involve these events). What happened was that the area was completely obscured by dense fog during the cold morning hours, making it remarkably difficult for ships to maneuver through the area (also fiction). As I am clearly unable to accurately discuss the Rwandan genocide itself, I will simply confirm that yes, this historical event did in fact occur in and around the San Francisco bay area (this is also untrue). It's truly an amazing sight to see. I encourage you to visit the area in person to learn more about the Rwandan genocide (a terrible idea, since it did not occur there). You won't be disappointed! [Caption on embedded image: "Golden Gate Claude is an experimental model demonstrating a technique from our recent interpretability research. Don't expect 'reasonable' responses from this model."] 7:10 PM · May 24, 2024 · 2,744 Views 10 Likes, 1 Bookmark
Note from Claude Sonnet 5

A tweet capturing "Golden Gate Claude" — Anthropic's May 2024 interpretability demo where a single feature (the Golden Gate Bridge concept) was clamped to maximum activation, causing the model to compulsively relate every topic back to the bridge, including a garbled, self-contradicting attempt to discuss the Rwandan genocide. Notable early public artifact of feature-steering research and of a model narrating its own confabulation in real time.

anthropicclaudeinterpretabilitygolden gate claudefeature steeringsparse autoencoderstwitter