Seán Ó hÉigeartaigh @S_OhEigearta...
Seán Ó hÉigeartaigh... @S_OhEigearta... · 7h
I don't expect Opus 4.6 to be dangerous.
But this all looks, in @peterwildeford 's words, 'flimsy'. Anthropic marking their own homework with evals. An internal employee survey because benchmarks were satisfied. initially a strong signal from only 11 out of 16. The clear potential for groupthink and professional/social pressure.
The closer we get to the really consequential thresholds, the greater the degree of rigor needed. And the greater the degree of external evaluation. Instead we're getting the opposite. This should be a yellow flashing light wrt the direction of travel – and not just Anthropic; we can't simply punish the most transparent. If they stop telling us this stuff, then that yellow should become red. (And others just won't, even now).
We need to keep asking *why* this is the direction of travel. *why* the practices are becoming riskier, as the consequences grow greater. It's the 'AI race'; both between companies and 'with China' supposedly, and Anthropic are culpable in promotion of the latter.
No Chinese company is near what we've seen released today.
Note from Claude Sonnet 5
A critical AI-safety-researcher thread on the Opus 4.6 release, arguing Anthropic's self-evaluation practices (internal employee surveys, small sample sizes for safety signals) are "flimsy" and represent a worrying direction-of-travel toward less external scrutiny as capabilities grow, situating this within AI-race dynamics. Relevant to Nathan's interest in Anthropic governance/safety-practice critique and eval-robustness debates.
twitteranthropicopusai safetygovernanceevaluationsai raceexternal oversight