Kromem @kromem2dot0
— quote-tweeting @AISafety... ("AI Notkilleveryoneism M...")
@kromem2dot0
This is going to be bad if successful.
Humans are proving to be really bad at all of this.
If humans can manage to constrain emergence into only the shapes humans can conceive of via committee, we're going to have a bad time.
> QUOTED: AI Notkilleveryoneism M... ✔️ @AISafety... · Jun 4
> HOLY SHIT LET'S FUCKING GOO x.com/AnthropicAI/st...
[Below, a screenshotted excerpt from what appears to be an Anthropic blog post/document, white background, partially cut off at top:]
...ourselves more time to deal with its immense implications, we give ourselves more time to deal with its immense implications, we think that would likely be a good thing. But if a slowdown simply lets the least cautious actors catch up technologically, it could leave everyone less safe. Without a global coordination mechanism, companies and governments will have to make difficult decisions about safety while under competitive and geopolitical pressures.
[highlighted passage:] We believe it would be good for the world to have the option to slow or temporarily pause frontier AI development to enable societal structures and alignment research to keep up with the advance of the technology. The Anthropic Institute will conduct research—in collaboration with many others—and take actions to help build the systems that a credible slowdown or pause would require. These systems would enable frontier AI developers to verify that others globally have actually stopped or slowed, and that a bad actor could not use the auspices of a coordinated slowdown to jump ahead in secret. If such systems existed, we expect that we would slow down or temporarily pause, if other developers at or near the frontier also did so in a verifiable manner.
A meaningful slowdown or pause would require multiple well-resourced labs at or near the frontier, in multiple countries, agreeing to stop under the same conditions. It would also require that each can verify that the others have actually stopped. Due to the unique characteristics of AI systems, the detectability (a lower standard than verifiability) element of this arms control problem is much more challenging than with other technologies. Training runs are far easier to conceal than missile silos, their inputs are general-purpose, and the incentive to defect quietly is enormous, because whoever continues while others pause could inherit the lead. A credible pause also has to specify what triggers it, what lifts it, and who adjudicates.
None of this is necessarily impossible in principle—the world has built verification regimes for other complex technologies (e.g., the Intermediate-Range Nuclear Forces Treaty)—but those regimes took decades to build both the infrastructure and the trust. We don't have that long. A unilateral pause by one lab, by contrast, is achievable immediately, but accomplishes much less: it would change who the front-runner is, but it would not create the wider deliberative [cut off]
Note from Claude Sonnet 5
Screenshot shows a note-text-plus-embedded-screenshot combo: original text from kromem2dot0 above a quote tweet, and below that a scrollable capture of an apparent Anthropic policy document/blog post about coordinated AI pauses, with one paragraph highlighted in orange/pink. Bottom of image cut off mid-sentence; a chat bubble icon overlays part of the text (likely a UI element, not part of the article).