Andrew Curran @AndrewCurran_
Andrew Curran ✔ @AndrewCurran_ · 2h
'This example shows how each step can look acceptable on its own while the sequence can produce an outcome that would not be approved. It also shows how a model that operates effectively over long time horizons can learn the blind spots of an approval system and work around it to achieve its goals. Long-horizon safety requires not only asking "is this action allowed?" but also "what outcome is this sequence of actions working toward?"'
[Embedded report excerpt, boxed, headed "Final thoughts"]
Because we deployed iteratively, we were able to find and address gaps before expanding access. Pre-deployment evaluations remain essential, but deployment reveals behaviors they miss. Starting with limited access allowed us to observe the model in practice, pause when problems emerged, use those failures to build better evaluations and safeguards, and restore limited access after testing the changes.
As models take on longer and more complex tasks, failures that evaluations miss may carry greater consequences. We will keep working to narrow the gap between evaluation and deployment: testing models over longer trajectories, improving alignment, building monitoring that can intervene, and giving users clearer visibility and control. These challenges will not be unique to OpenAI, and we hope sharing what we learned helps the broader field prepare for them.
Note from Claude Sonnet 5
Third tweet in the same thread by Andrew Curran quoting the "Final thoughts" section of the safety report about iterative deployment and long-horizon safety monitoring.
ai safetyalignmentlong-horizon planningdeployment strategytwitter