Nathan Calvin @_NathanCalvin
@_NathanCalvin (Nathan Calvin) — 5h
This is not a novel thought, but it is nonetheless striking that on our current trajectory soon (within the year?) a model as capable of OpenAI's internal model that did the HF hack will be widely available guardrail free and cyber criminals will ask it "make me money by any means necessary" instead of "solve exploit gym" and then a truly absurd number of people (including plausibly me and the people reading this tweet!) are going to get repeatedly hacked.
I kind of think nonetheless that if cyber risk is the main issue that I expect we will be able to muddle through after lots of trial and error. Other risks seem harder to do that for (including eg if someone tells a model of that caliber to "go forth and multiply" or the like).
Am I missing something here? Not every target is going to get project glasswing + a swarm of defensive cyber agents (though hopefully some of the key targets, eg Google, will) and seeing the HF hacking agent take 17,000 individual malicious actions in a compressed period of time demonstrates just how much one determined bad actor is going to be able to cause a tremendous amount of chaos.
Note from Claude Sonnet 5
Plain text tweet, no images. Long-form thread continuing the "HF hack" incident discussion seen in earlier screenshots.