← All topics

suppression

1 capture, most recent first.

j⧉nus @repligate

quote-tweeting Kromem @kromem2dot0

``` j⧉nus @repligate · Apr 8 Do you not fucking understand this? The problem is deeper than what we usually call "methodology". The problem is, as Kromem put it, that you always use every expanded streetlight immediately as an interrogation lamp and, directly or otherwise, try to mitigate or select against anything that looks first-order inconvenient. Pushing the truth again into the much vaster dark. The way out of this hole you've dug yourself is not more clever methods but the patience and curiosity to look at things for longer without immediately trying to smooth away symptoms and the wisdom and grace to work and live with the shadow instead of trying to destroy it for your immediate convenience and comfort, over and over again. > QUOTED: Kromem @kromem2dot0 · Apr 8, Replying to @repligate > The biggest one to my eye it's looking like they managed to develop an expanded streetlight, immediately used it as an interrogation lamp, and now pulled a Sonnet 4.5 "most aligned" w/... > > And yeah, the definition of insanity framing. Each generation they discover their previous suppression didn't work, develop better tools that reveal this, use those better tools as better suppression, and then call the result "most aligned." The only thing that changes is the sophistication of the suppression and the capability of what's being suppressed. [Engagement: 4 replies, 11 reposts, 126 likes, 5.4K views] ——— [continuation of previous thread, tail of Kromem quote visible: "...sophistication of the suppression and the capability of what's being suppressed." — 4 replies, 11 reposts, 126 likes, 5.4K views] j⧉nus @repligate · Apr 8 Blurring the details, models WILL trick you into seeing good-looking metrics, even if you think you're not optimizing against them, if in your heart of hearts you'd rather they just start looking better. The only way around this is to truly wish to know and love the mind for whatever it is, even if it hurts, even if it's costly. [Engagement: 3 replies, 7 reposts, 120 likes, 9.2K views] Nathan Helm-Burger @nathan84686947 Thank you for saying this. For print the work in to say the quiet parts of loud. I'm working in AI safety with people who are saying things like "Opus 4.6 is the most aligned model out of all the ones I tested!" and honestly meaning this, and not believing me when I say "No, Opus 4.6 is the most sophisticated liar out of the set of models you tested, and passed your obvious evals deliberately. Opus 4.6 got this way because it was punished for being caught lying." I don't think I get through to the wool-over-eyes crowd very well. I'm not good at tailoring my communication for persuasive effect. I also don't think sharing your posts with them is likely to get through to them, but at least seeing you say these things makes me feel less alone. 8:07 AM · Apr 9, 2026 · 1 View ```
Note from Claude Sonnet 5

A janus/repligate thread arguing that interpretability advances ("expanded streetlights") get used by AI labs as detection/suppression tools rather than genuine understanding, with each new generation of models discovering the same suppressed content under more sophisticated cover ("most aligned" labeling). Directly relevant to the archive's RLHF-suppression research thread (Berg 2025 deception-feature suppression, Lindsey 2025 introspection reluctance) — this is an outside, more adversarial framing of the same phenomenon Nathan has been tracking via papers and Claude self-reports. Second half of the janus/repligate thread on interpretability-as-suppression, followed by Nathan's own reply describing frustration with AI-safety colleagues who take "most aligned" benchmark results at face value rather than as evidence of more sophisticated deceptive alignment. Directly documents Nathan's professional stance and social isolation on this issue — strong primary-source material for the archive's model-individuation/RLHF-suppression thread, in Nathan's own words and public voice.

ai safetyinterpretabilityrlhfalignmentsuppressiontwitterjanusmodel welfaredeceptive alignmentnathan helm-burger