— saved image
Andrew Curran reposted elie @eliebakouch · 4h wow so seems like openai now has chain of thought monitoring at scale? wondering what technical challenges they faced here, would be something really cool to share more openly as this will contribute to making every model safer [Embedded image, titled "Steps we are taking":] Accordingly, we have scaled up robustness testing of our safeguards and security controls so that they are appropriate for a deployment of these capabilities. Internally, we have also taken the following steps so that further development of this model happens safely and securely: - We are implementing stricter security controls for higher-capability models and associated activities, including isolated testing environments, restricted network and tool access, enhanced model weight protections and encryption, additional monitoring and detection capabilities, and sandboxed execution. - We are pausing internal activities involving Astra that do not yet meet these strengthened security control requirements. - We have implemented universal monitoring for risky actions and misalignment across all agentic applications of Astra, including training and evaluation. [highlighted:] Monitors evaluate the model's Chain of Thought and trigger a security response to review and interrupt high risk activity. - We will work with relevant government agencies and select AI safety organizations to test the capabilities for this model. - We will be providing recommended security controls to third-party testing partners for running higher risk evaluations and workloads safely. elie @eliebakouch · Aug 6 one key piece of information here is that they don't have (any?) chain of thought monitoring for "rogue behavior" and also don't seem to have it for "gibberish text", otherwis... [cut off]
Note from Claude Sonnet 5
Twitter thread discussing an OpenAI safety document excerpt (re: a model codenamed 'Astra') describing security controls including chain-of-thought monitoring for risky/misaligned agentic behavior, with commentary from AI researcher Elie Bakouch questioning gaps in the monitoring.