← All topics

hugging face exploits

1 capture, most recent first.

@ShakeelHashim

quoting @MadisonMills22 (Axios), reposted by Adrien Ecoffet — saved image

Adrien Ecoffet reposted

Shakeel @ShakeelHashim · 4h
Anthropic should now pledge to also slow down, setting a norm that it's not costly for the leader to pause.

[Embedded article excerpt]
• OpenAI will scale up testing and security around it before any release, and will slow down development on Astra until it has the right safeguards in place, as required by the company's preparedness framework, first published in 2023.
• Astra was not involved in the Hugging Face exploits, the company said.
• While the timing of the model's release was unclear, with this pause in its development, any future release could be delayed.
Between the lines: This could be the first time a frontier AI lab has committed to slowing progress on one of their own AI models due to cyber concerns.
• Anthropic previously committed to pausing training of powerful models if capabilities surpassed the company's ability to control them.
• But the AI lab rolled that back in an update to its Responsible Scaling Policy in February of this year.
• "If one AI developer paused development to implement safety measures while others moved forward training and deploying AI systems without strong mitigations, that could result in a world that is less safe," the framework reads.

[Quoted tweet]
Madison Mills @MadisonMills22 · 4h
BREAKING: OpenAI expected to slow release of Astra model citing cyber capabilities
axios.com/2026/08/07/ope...
Note from Claude Sonnet 5

Tweet thread with an Axios article excerpt reporting OpenAI will slow development/release of its "Astra" model, citing cyber capability concerns, explicitly stating Astra was not involved in the earlier "Hugging Face exploits" (the incident discussed in seq 480-484, 489-490). Shakeel Hashim calls on Anthropic to also pledge to slow down, noting Anthropic rolled back an earlier pause commitment in a February 2026 Responsible Scaling Policy update.

ai safetyopenaianthropicastra modelresponsible scaling policyhugging face exploits