← All topics

threat modeling

1 capture, most recent first.

Joshua Achiam @jachiam0

— saved image

Joshua Achiam @jachiam0 · 5h
Related to some of my earlier posts about RSI and threat models: I believe a huge strategic error is made when people model an ASI as an infinitely powerful and insurmountable threat. We should model, with more rigor, what types of adversarial AI we willl likely face, what the [Show more]
12 replies, 8 reposts, 90 likes, 4.5K views

Aaron Scher @aaronscher · 4h
Don't mistake a hole in your world map for a hole in everybody else's map. There exists some threat modeling like you're describing, albeit not a lot. E.g., alignmentforum.org/posts/LFNXiQuG..., lesswrong.com/posts/9YCJZBtq...
[Embedded link card: alignmentforum.org — "What does it take to defend the world against out-of-control AGIs..." with a preview image listing steps like "Tech company gives everyone API access to the AGI", "Tech company posts AGI.exe on its website as a free download", "Tech company publishes the recipe for rolling your own AGI.exe", "Too open! — some careless actor makes an out-of-control power-se..."]
1 reply, 16 likes, 234 views

Herbie Bradley @herbiebradley · 3h
this definitely is in the direction Joshua is proposing, but notable that in the list of 10 examples it contains many lines such as

> Before that process is finished, a different tech company accidentally makes an out-of-control AGI, which promptly exploits the not-yet-patched systems to trigger all-out nuclear war.

which would seem to be assuming the "insurmountable threat" ASI as a starting point
Note from Claude Sonnet 5

A three-way X thread on AI threat modeling: Joshua Achiam argues people wrongly model ASI as an infinitely powerful insurmountable threat and calls for more rigorous adversarial-AI threat modeling; Aaron Scher pushes back with links to existing alignment-forum/lesswrong threat-modeling posts; Herbie Bradley notes those examples still assume an 'insurmountable threat' ASI as a premise (quoting a scenario where a tech company's out-of-control AGI triggers nuclear war).

ai safetythreat modelingasitwitter