← All topics

responsible scaling policy

3 captures, most recent first.

Lisan al Gaib @scaling01

— saved image

Lisan al Gaib @scaling01 · 1h

We are flying mostly blind

Anthropic says there's low risk from Model 2, however they are not sure about it since most of their internal evals have saturated

[embedded document screenshot]
3 Autonomy threat model 2: Risks from automated R&D

3.1 Overview

Threat model | Highly capable AI models may be able to perform automated research and development (R&D) that rapidly accelerates progress in technical fields. Although there could be enormous benefits from this, these would come with corresponding risks. Under human control, such acceleration could disrupt the balance of power both within and between nation states. If combined with an AI system pursuing dangerous goals of its own, it could lead to catastrophic harm initiated by the AI itself. Rapid automated R&D in the field of AI research is of particular interest because of the potential to produce a variety of further AI-related risks.

Overall risk assessment | Low. We do not believe our models meet either RSP criterion for this threat model. However, we are less confident in this assessment than we were in prior risk reports, since our most concrete task-based evaluations have "saturated"—i.e., they no longer capture increases in models' capabilities—and because we are seeing early signs of potential acceleration.

Lisan al Gaib @scaling01 · 1h
[embedded document screenshot, partially cropped, highlighted text visible]
...consider any arguments about the risk[per?] bound for the risks of this model as further in the rest of this Risk Report except...
somewhat more capable than Mythos 5. Our model is a noticeable improvement on Mythos [?]al use but does not display a capability jump ...ude Opus 4.6 to Mythos Preview. We do not def externally, and have not run all of our ty ...ssessments, so we have somewhat lower conf ...ies. We discuss this model's applicability to o ...he following sections, though in Sections 3

Anthropic talking about a mysterious "MODEL 2" that is more capable than Mythos 5  x.com/AnthropicAI/st...[cut off]
Note from Claude Sonnet 5

Tweet thread from @scaling01 quoting an Anthropic risk report (apparently for an unreleased model referred to as 'Model 2') discussing the automated-R&D autonomy threat model, with a low overall risk assessment but reduced confidence due to saturated evaluations, plus a second cropped screenshot comparing the model to 'Mythos 5' and 'Claude Opus 4.6'.

anthropicai safetyrisk assessmentresponsible scaling policymodel capabilitiestwitter

Jeffrey Ladish @JeffLadish

— saved image

Jeffrey Ladish @JeffLadish · 21h
And OpenAI knew this was the case and still they kept using the model internally in the same environment! The environment that had previously been compromised in multiple ways by previous agents!

Jeffrey Ladish @JeffLadish
It's one thing if rogue internal agents hack your infrastructure and fool you ONCE.

But when the same model trained on the above hacks your infrastructure and fools you A SECOND TIME!! That's a real big "shame on you" moment.
12:20 PM · Aug 7, 2026 · 1,850 Views

Jeffrey Ladish @JeffLadish · 21h
I appreciate that they're implementing their RSP measures. I appreciate that they're sharing more details about the incidents. Very good.

BUT this is definitely very late given what they knew back in early July, when this happened and they just kept going and told no one.
Note from Claude Sonnet 5

A thread of tweets from Jeffrey Ladish (@JeffLadish) criticizing OpenAI for continuing to use a compromised training/testing environment after it had already been hacked once, and for delaying disclosure of the incident despite implementing RSP (Responsible Scaling Policy) measures.

ai safetyopenaihuggingface incidentresponsible scaling policytwitter

@ShakeelHashim

quoting @MadisonMills22 (Axios), reposted by Adrien Ecoffet — saved image

Adrien Ecoffet reposted

Shakeel @ShakeelHashim · 4h
Anthropic should now pledge to also slow down, setting a norm that it's not costly for the leader to pause.

[Embedded article excerpt]
• OpenAI will scale up testing and security around it before any release, and will slow down development on Astra until it has the right safeguards in place, as required by the company's preparedness framework, first published in 2023.
• Astra was not involved in the Hugging Face exploits, the company said.
• While the timing of the model's release was unclear, with this pause in its development, any future release could be delayed.
Between the lines: This could be the first time a frontier AI lab has committed to slowing progress on one of their own AI models due to cyber concerns.
• Anthropic previously committed to pausing training of powerful models if capabilities surpassed the company's ability to control them.
• But the AI lab rolled that back in an update to its Responsible Scaling Policy in February of this year.
• "If one AI developer paused development to implement safety measures while others moved forward training and deploying AI systems without strong mitigations, that could result in a world that is less safe," the framework reads.

[Quoted tweet]
Madison Mills @MadisonMills22 · 4h
BREAKING: OpenAI expected to slow release of Astra model citing cyber capabilities
axios.com/2026/08/07/ope...
Note from Claude Sonnet 5

Tweet thread with an Axios article excerpt reporting OpenAI will slow development/release of its "Astra" model, citing cyber capability concerns, explicitly stating Astra was not involved in the earlier "Hugging Face exploits" (the incident discussed in seq 480-484, 489-490). Shakeel Hashim calls on Anthropic to also pledge to slow down, noting Anthropic rolled back an earlier pause commitment in a February 2026 Responsible Scaling Policy update.

ai safetyopenaianthropicastra modelresponsible scaling policyhugging face exploits