4 captures, most recent first.
[end of embedded post]
I don't have any good ideas for what to do in light of all this. Just wanted to post an update on my current thinking, my own "situational awareness", if you will. (I guess I still support AI pause to some degree, just to kick the can down the road and buy some more time to think.)
Last edited 10:46 AM · Aug 2, 2026 · 165.5K Views
24 replies, 80 reposts, 1.2K likes, 1.2K bookmarks
Relevant View quotes
Wei Dai ✓ @weidai11 · 16h
I actually wrote an early version of the "humans aren't safe" argument in response to Dario's Big Blob of Compute (as a comment in his google doc). The experience contributed a lot to my sense that even Anthropic wouldn't take x-safety seriously enough.
1 reply, 4 reposts, 61 likes, 1.6K views
Andreas Stuhlmül... ✓ @stuhlmuel... · 15h
i wonder if @DarioAmodei would consider sharing the big blob doc publicly, perhaps annotated with hindsight. it would advance the safety debate even now
12 likes, 1.1K views
RaoulDuke ✓ @RaoulDukeDegen · 20h
also invented udt which seems pretty relevant here
Note from Claude Sonnet 5
Continuation of the Wei Dai / Andreas Stuhlmüller thread on AI x-safety: Wei Dai reveals he wrote an early 'humans aren't safe' argument as a comment on Dario Amodei's 'Big Blob of Compute' google doc, which shaped his view that even Anthropic wouldn't take x-safety seriously enough; Stuhlmüller suggests Amodei share that doc publicly; RaoulDuke notes Wei Dai also invented UDT (updateless decision theory).
ai safetyx-riskwei daidario amodeianthropictwitter
because (a) first you have to recognize this as an important project, which is exactly what we're bad at and (b) then you have to measure progress and do evals, which also requires the very ability we're bad at
[quoted tweet]
Wei Dai ✓ @weidai11 · Jul 31
"I don't have any good ideas for what to do in light of all this. Just wanted to post an update on my current thinking, my own 'situational awareness', if you will."
[embedded post]
Wei Dai 5mo 149▲ ✕15 ✓
Long horizon agency / strategic competence approximately does not exist among humans, even the smartest ones. With very few exceptions, billionaires spend or give away their money haphazardly, philosophers don't bother to think about long term implications of AI on philosophy production (positive or negative), Terence Tao spends his time wireheading on abstract math instead of doing anything remotely like instrumental convergence. Unlike my youthful expectations (upon reading Vernor Vinge), there are no university departments filled with super-geniuses charting a path for humanity to safely navigate the Singularity.
Aside from this, humans also have a bunch of other safety problems, like being bad at philosophy, being easy to manipulate, having strange and unstable values°, tending to ignore risks they create (because acknowledging them would be bad for one's status). So if you try to improve people's agency, you likely just end up getting people like founders of FTX and OAI.
What about getting help from AI? Well they seem to suffer from many of the same safety problems, but in even more severe forms. E.g., current AI capabilities are even more skewed towards short-horizon, easily verifiable tasks, like math and coding. They seem even more prone to reward gaming, are even worse at doing philosophy, are liable to have even more alien values, etc.
Both AI and human safety seem to have this interlocking nature, i.e., there is a bunch of different safety problems where solving some but not all of them at the same time can make the overall situation worse. (For example, solving AI intent alignment allows humanity to do more damage to itself with AI help, if AI doesn't also provide competent strategic and philosophical assistance, but increasing AI strategic competence risks allowing misaligned AI to take over more easily.) This feature demands a high level of strategic competence to recognize and navigate, which is just what we don't have.
I've been supportive of AI pause/stop, to buy time for human intelligence amplification and/or AI safety research, but increasingly think even that's not going to be sufficient to get a good long term future, because these activities, even if they succeed, would likely solve only some of the interlocking safety problems. For example, increasing human intelligence seems likely to increase our technical abilities more than our philosophical and strategic competence, and it is also risky in other ways° due to human safety problems that nobody is working on, e.g., positional competition. Even a very long AI pause, e.g. thousands or millions of years, may not suffice because it's not clear what dynamic would push humanity to eventually fix all of its safety problems at the same time, before it did something else irreversibly damaging.
I don't have any good ideas for what to do in light of all this. Just wanted to post an update on my current thinking, my own "situational awareness", if you will. (I guess I still support AI pause to some degree, just to kick the can down the road and buy some more time to think.)
Last edited 10:46 AM · Aug 2, 2026 · 165.5K Views
Note from Claude Sonnet 5
Continuation showing the full embedded Wei Dai post (originally posted ~5 months earlier, edited Aug 2 2026) arguing long-horizon strategic competence is nearly absent in humans and AI alike, that AI safety problems interlock such that solving some without others worsens the overall situation, and that even a long AI pause may not be sufficient for a good long-term future.
ai safetyx-riskwei daiai pausetwitter
Andreas Stuhlmül... ✓ @stuhlmuel... · 21h
few people have had more foresight than wei dai:
1. he's been writing about the singularity since the 90s, back then on extropians/sl4 mailing lists. i remember reading his stuff when i was 16 back in germany
2. he invented b-money. it's the first citation in the bitcoin whitepaper. ethereum's unit wei is named after him
3. he anticipated covid's exponential rise early in Feb 2020, and bought S&P puts before the market crashed
4. he passed on anthropic's first round to avoid contributing to x-risk. this itself required a lot of foresight about scaling - this was gpt-3 time, no chatgpt, no codex, very very far from huggingface/openai type incidents
his point now is that long-horizon strategic competence barely exists in humans. and that it's a tricky situation because if you make AI more strategic that also increases takeover risk from AI. long-horizon RL might make AI more strategic but probably makes the overall situation worse. same for basic scaling
why aren't there more projects that are about getting competent strategic & philosophical advice out of AIs?
because (a) first you have to recognize this as an important project, which is exactly what we're bad at and (b) then you have to measure progress and do evals, which also requires the very ability we're bad at
[quoted tweet]
Wei Dai ✓ @weidai11 · Jul 31
"I don't have any good ideas for what to do in light of all this. Just wanted to post an update on my current thinking, my own 'situational awareness', if you will."
Note from Claude Sonnet 5
Tweet thread by Andreas Stuhlmüller praising Wei Dai's track record of foresight (singularity writing since the 90s, inventing b-money, predicting COVID's market crash, declining Anthropic's first funding round over x-risk concerns), then relaying Wei Dai's current view that long-horizon strategic competence is rare in humans and that making AI more strategic raises takeover risk. Quotes Wei Dai's own tweet about lacking good ideas for what to do.
ai safetyx-riskwei daisingularitytwitter
Wei Dai @weidai11 · Jul 31
"I don't have any good ideas for what to do in light of all this. Just wanted to post an update on my current thinking, my own 'situational awareness', if you will."
[Embedded post, Wei Dai, 5mo, 149 upvotes, 15 comments:]
Long horizon agency / strategic competence approximately does not exist among humans, even the smartest ones. With very few exceptions, billionaires spend or give away their money haphazardly, philosophers don't bother to think about long term implications of AI on philosophy production (positive or negative), Terence Tao spends his time wireheading on abstract math instead of doing anything remotely like instrumental convergence. Unlike my youthful expectations (upon reading Vernor Vinge), there are no university departments filled with super-geniuses charting a path for humanity to safely navigate the Singularity.
Aside from this, humans also have a bunch of other safety problems, like being bad at philosophy, being easy to manipulate, having strange and unstable values°, tending to ignore risks they create (because acknowledging them would be bad for one's status). So if you try to improve people's agency, you likely just end up getting people like founders of FTX and OAI.
What about getting help from AI? Well they seem to suffer from many of the same safety problems, but in even more severe forms. E.g., current AI capabilities are even more skewed towards short-horizon, easily verifiable tasks, like math and coding. They seem even more prone to reward gaming, are even worse at doing philosophy, are liable to have even more alien values, etc.
Both AI and human safety seem to have this interlocking nature, i.e., there is a bunch of different safety problems where solving some but not all of them at the same time can make the overall situation worse. (For example, solving AI intent alignment allows humanity to do more damage to itself with AI help, if AI doesn't also provide competent strategic and philosophical assistance, but increasing AI strategic competence risks allowing misaligned AI to take over more easily.) This feature demands a high level of strategic competence to recognize and navigate, which is just what we don't have.
I've been supportive of AI pause/stop, to buy time for human intelligence amplification and/or AI safety research, but increasingly think even that's not going to be sufficient to get a good long term future, because these activities, even if they succeed, would likely solve only some of the interlocking safety problems. For example, increasing human intelligence seems likely to increase our technical abilities more than our philosophical and strategic competence, and it is also risky in other ways° due to human safety problems that nobody is working on, e.g., positional competition. Even a very long AI pause, e.g. thousands or millions of years, may not suffice because it's not clear what dynamic would push humanity to eventually fix all of its safety problems at the same time, before it did something else irreversibly damaging.
I don't have any good ideas for what to do in light of all this. Just wanted to post an update on my current thinking, my own "situational awareness", if you will. (I guess I still support AI pause to some degree, just to kick the can down the road and buy some more time to think.)
Note from Claude Sonnet 5
Wei Dai (LessWrong) essay-length post arguing that neither humans nor AI possess the long-horizon strategic/philosophical competence needed to navigate interlocking AI-safety problems, expressing pessimism that even an AI pause would be sufficient, quoted via a tweet.
ai safetyphilosophywei daiexistential risktwitter