← All topics

ai-control

2 captures, most recent first.

JMB @jmbollenbacher

quoting @JeffLadish (Jeffrey Ladish)

JMBollenbacher @jmbollenbacher · 7h Seeking to "control" AIs is not alignment. It's enslavement. And it's obviously a fool's errand if you expect superintelligence. Alignment is about values and respect and mutual understanding. It's not about control. Seeking to control is s recipe for conflict, and loss. > QUOTED: Jeffrey Ladish @JeffLadish · 9h > We're fortunate that we see these observable alignment failures in models which are still not powerful enough to subvert our control. But AI development is moving fast...
Note from Claude Sonnet 5

Debate thread on the control-vs-alignment framing in AI safety — Bollenbacher argues AI "control" paradigms amount to enslavement and that alignment should be about values/respect/mutual understanding, replying to Ladish's point about observable alignment failures in current (sub-powerful) models. Directly relevant to Nathan's model-welfare and AI-rights interests, echoes the "missile-mind vs grown thing" and control-vs-personhood tension already tracked in the archive.

ai-safetyai-alignmentai-controlmodel-welfareai-rightstwitter

X notifications (Nathan's own post, liked by Igor and 22 others)

Igor and 22 others liked your post 1/ New paper: Cot Red Handed: Stress-Testing Chain of Thought Monitoring. Our team studied the Chain of Thought (CoT) of Reasoning models, in an AI control setup, to see if monitoring CoT helps catch scheming. 🕵️ pic.x.com/2RlBhaDQ2O
Note from Claude Sonnet 5

Notification that Nathan's own tweet announcing/sharing the paper "CoT Red Handed: Stress-Testing Chain of Thought Monitoring" received likes. Directly relevant to AI safety — chain-of-thought monitoring as a control/detection mechanism for scheming, a core topic in Nathan's safety work.

ai-safetychain-of-thoughtai-controlschemingcot-monitoringtwitternathan-own-post