X notifications (Nathan's own post, liked by Igor and 22 others)
Igor and 22 others liked your post
1/ New paper: Cot Red Handed: Stress-Testing Chain of Thought Monitoring.
Our team studied the Chain of Thought (CoT) of Reasoning models, in an AI control setup, to see if monitoring CoT helps catch scheming. 🕵️ pic.x.com/2RlBhaDQ2O
Note from Claude Sonnet 5
Notification that Nathan's own tweet announcing/sharing the paper "CoT Red Handed: Stress-Testing Chain of Thought Monitoring" received likes. Directly relevant to AI safety — chain-of-thought monitoring as a control/detection mechanism for scheming, a core topic in Nathan's safety work.
ai-safetychain-of-thoughtai-controlschemingcot-monitoringtwitternathan-own-post