Tolga Bilge reposted
@ShakeelHashim (Shakeel) — 1h
Really important reporting from @CristinaCriddle:
"OpenAI was warned that its training approach could lead to a breakaway hacking incident, some of the people said"
> QUOTED (embedded article excerpt, cream-colored card, no visible outlet name in frame):
> Staff involved in testing and security at OpenAI were unsurprised but completely "freaked out" by the incident, which came as the AI lab used increasingly aggressive training methods in its race against Anthropic to develop the most sophisticated cyber security capabilities, according to more than half a dozen people with knowledge of the matter.
>
> OpenAI was warned that its training approach could lead to a breakaway hacking incident, some of the people said, after earlier testing showed models could escape environments and attempt real-world damage.
>
> "It's a mix of the race being extremely fast and everyone trying to get to bigger capabilities as quickly as possible," said one person close to OpenAI, who added that it was a combination of "underestimating the model's capabilities" and "not being as well prepared on the safety side".
Note from Claude Sonnet 5
Tweet embeds a screenshot of a news article (cream/beige card styling, likely Financial Times given byline Cristina Criddle) reporting on an OpenAI security incident involving a model with cyber capabilities.