← All topics

chain of thought monitoring

3 captures, most recent first.

@eliebakouch

— saved image

Andrew Curran reposted
elie @eliebakouch · 4h
wow so seems like openai now has chain of thought monitoring at scale? wondering what technical challenges they faced here, would be something really cool to share more openly as this will contribute to making every model safer

[Embedded image, titled "Steps we are taking":]
Accordingly, we have scaled up robustness testing of our safeguards and security controls so that they are appropriate for a deployment of these capabilities. Internally, we have also taken the following steps so that further development of this model happens safely and securely:
- We are implementing stricter security controls for higher-capability models and associated activities, including isolated testing environments, restricted network and tool access, enhanced model weight protections and encryption, additional monitoring and detection capabilities, and sandboxed execution.
- We are pausing internal activities involving Astra that do not yet meet these strengthened security control requirements.
- We have implemented universal monitoring for risky actions and misalignment across all agentic applications of Astra, including training and evaluation. [highlighted:] Monitors evaluate the model's Chain of Thought and trigger a security response to review and interrupt high risk activity.
- We will work with relevant government agencies and select AI safety organizations to test the capabilities for this model.
- We will be providing recommended security controls to third-party testing partners for running higher risk evaluations and workloads safely.

elie @eliebakouch · Aug 6
one key piece of information here is that they don't have (any?) chain of thought monitoring for "rogue behavior" and also don't seem to have it for "gibberish text", otherwis... [cut off]
Note from Claude Sonnet 5

Twitter thread discussing an OpenAI safety document excerpt (re: a model codenamed 'Astra') describing security controls including chain-of-thought monitoring for risky/misaligned agentic behavior, with commentary from AI researcher Elie Bakouch questioning gaps in the monitoring.

ai safetyopenaichain of thought monitoringsecurity controls

Samuel Hammond @hamandcheese

— saved image

Samuel Hammon... @hamandche... · 2h
I'm confused. In 2025, OpenAI made a public commitment to not optimize CoT and to monitor CoT for reward hacking.

Did they just ignore those commitments?

openai.com/index/chain-of...

[quoted OpenAI page]
OpenAI
Monitoring frontier reasoning models for reward hacking

We believe that CoT monitoring may be one of few tools we will have to oversee superhuman models of the future.

We have further found that directly optimizing the CoT to adhere to specific criteria (e.g. to not think about reward hacking) may boost performance in the short run; however, it does not eliminate all misbehavior and can cause a model to hide its intent. We hope future research will find ways to directly optimize CoTs without this drawback, but until then

We recommend against applying strong optimization pressure directly to the CoTs of frontier reasoning models, leaving CoTs unrestricted for monitoring.

[reply]
elie @eliebakouch · 22h
one key piece of information here is that they don't have (any?) chain of thought monitoring for "rogue behavior" and also don't seem to have it for "gibberish text", otherwis... [cut off]
Note from Claude Sonnet 5

Tweet from Samuel Hammond questioning whether OpenAI ignored its own 2025 public commitment to leave chain-of-thought (CoT) unoptimized and monitor it for reward hacking, quoting OpenAI's own blog post on CoT monitoring, with a reply from elie (@eliebakouch) noting OpenAI apparently lacks CoT monitoring for 'rogue behavior' or 'gibberish text'.

ai safetychain of thought monitoringopenaireward hackinginterpretability

Harlan Stewart @HumanHarlan

quoting @OpenAI, reposted by Bogdan Ionut Cirstea — saved image

Bogdan Ionut Cirstea reposted

Harlan Stewart @HumanHarlan · 22h
It's crazy that they said all that stuff about the importance of CoT monitoring and then didn't actually do it

[Quoted tweet]
OpenAI @OpenAI · Jul 15, 2025
Chain of Thought (CoT) monitoring could be a powerful tool for overseeing future AI systems—especially as they become more agentic. That's why we're backing a new research paper from a cross-institutional team of researchers pushin...[cut off]
Note from Claude Sonnet 5

Continuation of the same discussion thread about the "HF incident" (see seq 480-482): Harlan Stewart criticizes OpenAI for publicly emphasizing chain-of-thought monitoring in a July 2025 tweet but apparently not applying it in practice, implied to be relevant to how the incident went undetected.

ai safetychain of thought monitoringopenaialignment