Dimitris Papailiopoulos @DimitrisPapail
— reply from @PandaAshwinee (Ashwinee Panda)
Note from Claude Sonnet 5
Technical AI-research discussion: Claude Opus 4.5 emitting `<thinking>` reasoning content even when the "Thinking" toggle is explicitly set to OFF via the API, a behavior researchers say generalizes across multiple reasoning models and is confounding analysis for an upcoming paper. Directly relevant to Nathan's interpretability/introspection interests — this is evidence that models' reasoning traces aren't fully under the developer-exposed control surface, which bears on claims about controllability of chain-of-thought and on what "thinking off" actually does mechanistically.
twitterclaude opus 4.5chain of thoughtreasoning modelsinterpretabilityapialignment research