Peter Barnett @peterbarnett_ · 3h:
"This is crazy, and I think totally against the spirit of the original RSP. If Anthropic were sticking to its original commitments, this would probably require them to temporarily halt their AI development. (I expect the same goes for OpenAI)"
> QUOTED: Nathan Calvin @_NathanCalvin · 3h:
"To determine whether Opus 4.6 is ASL-4 on autonomous AI R&D, Anthropic did a survey of 16 employees b/c their benchmarks are saturated.
..."
[Embedded screenshot of Anthropic document text]:
"Similarly to Claude Opus 4.5, the ASL determination for autonomous AI R&D risks required careful judgment. Opus 4.6 has roughly reached the pre-defined thresholds we set for straightforward ASL-4 rule-out based on benchmark tasks. Thus the rule-out in this case is primarily informed by qualitative impressions of model capabilities for complex, long-horizon tasks and the results of a survey of Anthropic employees (described in Section 8.3.1 below). We may build more tasks to test capabilities in the long-horizon regime, but evaluations of model capabilities may also need to depend more on expert judgments, as they already do for other risks via expert red-teaming and uplift trials."
Note from Claude Sonnet 5
Substantive AI safety governance critique: Anthropic's ASL-4 (autonomous AI R&D) rule-out determination for Claude Opus 4.6 relied on a qualitative survey of 16 employees because benchmarks had saturated, which Peter Barnett argues violates the spirit of Anthropic's original Responsible Scaling Policy (RSP) commitments — implying a halt should have been triggered instead. Directly relevant to Nathan's AI governance/safety interests and to tracking Anthropic's own safety-commitment fidelity, a recurring archive theme (parallel to the Fable classifier incident already logged in memory).