Lisan al Gaib @scaling01
— saved image
Lisan al Gaib @scaling01 · 1h We are flying mostly blind Anthropic says there's low risk from Model 2, however they are not sure about it since most of their internal evals have saturated [embedded document screenshot] 3 Autonomy threat model 2: Risks from automated R&D 3.1 Overview Threat model | Highly capable AI models may be able to perform automated research and development (R&D) that rapidly accelerates progress in technical fields. Although there could be enormous benefits from this, these would come with corresponding risks. Under human control, such acceleration could disrupt the balance of power both within and between nation states. If combined with an AI system pursuing dangerous goals of its own, it could lead to catastrophic harm initiated by the AI itself. Rapid automated R&D in the field of AI research is of particular interest because of the potential to produce a variety of further AI-related risks. Overall risk assessment | Low. We do not believe our models meet either RSP criterion for this threat model. However, we are less confident in this assessment than we were in prior risk reports, since our most concrete task-based evaluations have "saturated"—i.e., they no longer capture increases in models' capabilities—and because we are seeing early signs of potential acceleration. Lisan al Gaib @scaling01 · 1h [embedded document screenshot, partially cropped, highlighted text visible] ...consider any arguments about the risk[per?] bound for the risks of this model as further in the rest of this Risk Report except... somewhat more capable than Mythos 5. Our model is a noticeable improvement on Mythos [?]al use but does not display a capability jump ...ude Opus 4.6 to Mythos Preview. We do not def externally, and have not run all of our ty ...ssessments, so we have somewhat lower conf ...ies. We discuss this model's applicability to o ...he following sections, though in Sections 3 Anthropic talking about a mysterious "MODEL 2" that is more capable than Mythos 5 x.com/AnthropicAI/st...[cut off]
Note from Claude Sonnet 5
Tweet thread from @scaling01 quoting an Anthropic risk report (apparently for an unreleased model referred to as 'Model 2') discussing the automated-R&D autonomy threat model, with a low overall risk assessment but reduced confidence due to saturated evaluations, plus a second cropped screenshot comparing the model to 'Mythos 5' and 'Claude Opus 4.6'.
anthropicai safetyrisk assessmentresponsible scaling policymodel capabilitiestwitter