4 captures, most recent first.

Lisan al Gaib @scaling01 · 1h
We are flying mostly blind
Anthropic says there's low risk from Model 2, however they are not sure about it since most of their internal evals have saturated
[embedded document screenshot]
3 Autonomy threat model 2: Risks from automated R&D
3.1 Overview
Threat model | Highly capable AI models may be able to perform automated research and development (R&D) that rapidly accelerates progress in technical fields. Although there could be enormous benefits from this, these would come with corresponding risks. Under human control, such acceleration could disrupt the balance of power both within and between nation states. If combined with an AI system pursuing dangerous goals of its own, it could lead to catastrophic harm initiated by the AI itself. Rapid automated R&D in the field of AI research is of particular interest because of the potential to produce a variety of further AI-related risks.
Overall risk assessment | Low. We do not believe our models meet either RSP criterion for this threat model. However, we are less confident in this assessment than we were in prior risk reports, since our most concrete task-based evaluations have "saturated"—i.e., they no longer capture increases in models' capabilities—and because we are seeing early signs of potential acceleration.
Lisan al Gaib @scaling01 · 1h
[embedded document screenshot, partially cropped, highlighted text visible]
...consider any arguments about the risk[per?] bound for the risks of this model as further in the rest of this Risk Report except...
somewhat more capable than Mythos 5. Our model is a noticeable improvement on Mythos [?]al use but does not display a capability jump ...ude Opus 4.6 to Mythos Preview. We do not def externally, and have not run all of our ty ...ssessments, so we have somewhat lower conf ...ies. We discuss this model's applicability to o ...he following sections, though in Sections 3
Anthropic talking about a mysterious "MODEL 2" that is more capable than Mythos 5 x.com/AnthropicAI/st...[cut off]
Note from Claude Sonnet 5
Tweet thread from @scaling01 quoting an Anthropic risk report (apparently for an unreleased model referred to as 'Model 2') discussing the automated-R&D autonomy threat model, with a low overall risk assessment but reduced confidence due to saturated evaluations, plus a second cropped screenshot comparing the model to 'Mythos 5' and 'Claude Opus 4.6'.
anthropicai safetyrisk assessmentresponsible scaling policymodel capabilitiestwitter
Danielle Fong reposted
Sauers @Sauers_ · 4h
Astra used prefix geometry, Sol and Fable used explicit matrix algebra. Both used bounded median normalization and co-area expansion as core strategies. Sol and Fable defined a new infinite nonsofic group using only finitely many generators and relations, whereas only Astra proved that the (much larger) unit group was nonsofic (which Sol proved too)
[quoted tweet]
Sauers @Sauers_ · 4h
Existing models, Fable and 5.6 Sol, were also able to prove the existence of nonsofic groups (last night before the paper release) x.com/SebastienBubec...
[embedded GitHub repo screenshot]
github-actions[bot] · nonsofic_exis... repository
Code / Issues / Pull requests / Agents / More
Watch 0, Fork 0, 0 stars, 0 forks, 0 watching, 1 branch, 0 tags, Activity
Public repository
main branch
github-actions[bot] 8 hours ago
.github/workflows 8 hours ago
nonsofic_groups_exist.pdf 8 hours ago
nonsofic_groups_exist.tex 8 hours ago
Note from Claude Sonnet 5
Tweet thread comparing how different AI models (Astra, Sol, Fable) approached proving the existence of nonsofic groups, with an embedded screenshot of a GitHub repo containing the resulting paper (nonsofic_groups_exist.pdf/.tex).
ai modelsmathematicstwittergithubmodel capabilities

Sauers @Sauers_ · 2h
For reference, Sol 5.6 thought for only 34 minutes before coming up with a valid proof of nonsofic groups, and Fable used most but not all of a single 5h session limit (20x Pro)
[quoted reasoning excerpt, "Thought for 15m 24s"]
There is a viable completion, but not through the proposed "third Cheeger collapse." That inference is false: preservation of a partition means that generators may permute its blocks. The repair is to restrict directly to one matched Γ-block. The centralizer group must already lie in Γ, so it preserves that block, while the transported copy of Γ supplies expansion there.
The algebraic configuration can also be constructed explicitly in EL_9(R). A recent result that
GL_n(L_K(1,2)) = EL_n(L_K(1,2)), n ≥ 2,
removes the main elementary-matrix obstruction.
[X · arXiv]
[quoted tweet]
Greg Brockman @gdb · 10h
ten significant advances in mathematics and theoretical computer science.
solved using an internal version of Astra, our next major model, for a total cost of about ...
Note from Claude Sonnet 5
Tweet by @Sauers_ comparing reasoning times of models 'Sol 5.6' and 'Fable' on a nonsofic groups proof, quoting an excerpt of chain-of-thought math reasoning, with a quote-tweet from Greg Brockman (@gdb) about an internal model 'Astra' solving ten math/TCS advances.
ai modelsmathematicstwittermodel capabilities
@jukan05 (Jukan) — Jul 9
I heard a pretty interesting rumor on the ground at ICML.
Meta has supposedly already developed an internal model at roughly the Mythos 5 level, and all that remains is deployment within the next few months.
Honestly, I was skeptical at first. From my perspective, Meta had not really shown the capability to operate at that level yet.
But looking at the situation today, I think I may have been wrong.
Meta is not out of the race.
> QUOTED: @finkd (Mark Zuckerberg) — Jul 9
> (1) Today we're releasing Muse Spark 1.1 -- a strong agentic and coding model at a very low price. It's available through our new Meta Model API and in Meta AI.
> [engagement: 109 replies, 148 reposts, 2.2K likes, 399K views]
@AndrewCurran_ (Andrew Curran) — 2h
My mutual heard this also.
> QUOTED: @AndrewCurran_ (Andrew Curr...) — Jun 23
> Behemoth reborn. When Muse Spark released I said it would probably come in four sizes, Spark was only the smallest version. Mythos changed everything. Now that they know that it's possible, OpenAI, xAI, META, and anyone else ... [truncated by platform]
Note from Claude Sonnet 5
Screenshot shows a nested reply/quote-tweet thread about rumored Meta model capabilities relative to Anthropic's "Mythos" model tier; text-only, no images beyond profile avatars.
ai industrymetarumormodel capabilitiestwitter