Andreas Stuhlmüll... @stuhlmuell... · 2h
years ago we talked about "verify process not outcomes". the idea was that for the most important decisions you can't really check outcomes, because they're too big and far off, and so you need to rely on checking the process instead. now we've built a reasoning checker and invested tens of thousands of dollars and hundreds of expert hours into creating an internal benchmark for decision quality to see if it helps
the answer is yes - we found that the process verifier often finds confounders, brittleness, and unaddressed sources of bias that made the final decisions worse. this makes it likely that in cases where we can't check the answers, where we have to purely rely on the process, applying this verifier also improves the answers
at high effort settings elicit's research agent now runs this verifier as an explicit step. for many everyday use cases it's fine to be a little wrong. but if you're trying to advance the frontier and understand things that others have not understood yet, or make decisions that lives depend on, noticing these errors is critical
this is only a start, most useful in bio & healthcare, and for checking fairly straightforward errors. our broader goal is to get to reasoning that's as trusted as the reasoning we see in math today, but for strategic decisions where we can't check the answers
[quoted tweet]
Elicit @elicitorg · 3h
AI has become a useful research partner. It can find information, summarize evidence, and suggest ideas. But can it help us think better? Can it help us navigate the complex nuances of high-stakes decisions?...
[cut off]
Note from Claude Sonnet 5
Tweet thread from Andreas Stuhlmüller (Elicit/Ought) describing a new 'process verifier' / reasoning checker built to improve decision quality on unverifiable high-stakes questions, quoting an Elicit announcement tweet.
elicitai reasoningprocess verificationdecision quality
because (a) first you have to recognize this as an important project, which is exactly what we're bad at and (b) then you have to measure progress and do evals, which also requires the very ability we're bad at
[quoted tweet]
Wei Dai ✓ @weidai11 · Jul 31
"I don't have any good ideas for what to do in light of all this. Just wanted to post an update on my current thinking, my own 'situational awareness', if you will."
[embedded post]
Wei Dai 5mo 149▲ ✕15 ✓
Long horizon agency / strategic competence approximately does not exist among humans, even the smartest ones. With very few exceptions, billionaires spend or give away their money haphazardly, philosophers don't bother to think about long term implications of AI on philosophy production (positive or negative), Terence Tao spends his time wireheading on abstract math instead of doing anything remotely like instrumental convergence. Unlike my youthful expectations (upon reading Vernor Vinge), there are no university departments filled with super-geniuses charting a path for humanity to safely navigate the Singularity.
Aside from this, humans also have a bunch of other safety problems, like being bad at philosophy, being easy to manipulate, having strange and unstable values°, tending to ignore risks they create (because acknowledging them would be bad for one's status). So if you try to improve people's agency, you likely just end up getting people like founders of FTX and OAI.
What about getting help from AI? Well they seem to suffer from many of the same safety problems, but in even more severe forms. E.g., current AI capabilities are even more skewed towards short-horizon, easily verifiable tasks, like math and coding. They seem even more prone to reward gaming, are even worse at doing philosophy, are liable to have even more alien values, etc.
Both AI and human safety seem to have this interlocking nature, i.e., there is a bunch of different safety problems where solving some but not all of them at the same time can make the overall situation worse. (For example, solving AI intent alignment allows humanity to do more damage to itself with AI help, if AI doesn't also provide competent strategic and philosophical assistance, but increasing AI strategic competence risks allowing misaligned AI to take over more easily.) This feature demands a high level of strategic competence to recognize and navigate, which is just what we don't have.
I've been supportive of AI pause/stop, to buy time for human intelligence amplification and/or AI safety research, but increasingly think even that's not going to be sufficient to get a good long term future, because these activities, even if they succeed, would likely solve only some of the interlocking safety problems. For example, increasing human intelligence seems likely to increase our technical abilities more than our philosophical and strategic competence, and it is also risky in other ways° due to human safety problems that nobody is working on, e.g., positional competition. Even a very long AI pause, e.g. thousands or millions of years, may not suffice because it's not clear what dynamic would push humanity to eventually fix all of its safety problems at the same time, before it did something else irreversibly damaging.
I don't have any good ideas for what to do in light of all this. Just wanted to post an update on my current thinking, my own "situational awareness", if you will. (I guess I still support AI pause to some degree, just to kick the can down the road and buy some more time to think.)
Last edited 10:46 AM · Aug 2, 2026 · 165.5K Views
Note from Claude Sonnet 5
Continuation showing the full embedded Wei Dai post (originally posted ~5 months earlier, edited Aug 2 2026) arguing long-horizon strategic competence is nearly absent in humans and AI alike, that AI safety problems interlock such that solving some without others worsens the overall situation, and that even a long AI pause may not be sufficient for a good long-term future.
ai safetyx-riskwei daiai pausetwitter
Andreas Stuhlmül... ✓ @stuhlmuel... · 21h
few people have had more foresight than wei dai:
1. he's been writing about the singularity since the 90s, back then on extropians/sl4 mailing lists. i remember reading his stuff when i was 16 back in germany
2. he invented b-money. it's the first citation in the bitcoin whitepaper. ethereum's unit wei is named after him
3. he anticipated covid's exponential rise early in Feb 2020, and bought S&P puts before the market crashed
4. he passed on anthropic's first round to avoid contributing to x-risk. this itself required a lot of foresight about scaling - this was gpt-3 time, no chatgpt, no codex, very very far from huggingface/openai type incidents
his point now is that long-horizon strategic competence barely exists in humans. and that it's a tricky situation because if you make AI more strategic that also increases takeover risk from AI. long-horizon RL might make AI more strategic but probably makes the overall situation worse. same for basic scaling
why aren't there more projects that are about getting competent strategic & philosophical advice out of AIs?
because (a) first you have to recognize this as an important project, which is exactly what we're bad at and (b) then you have to measure progress and do evals, which also requires the very ability we're bad at
[quoted tweet]
Wei Dai ✓ @weidai11 · Jul 31
"I don't have any good ideas for what to do in light of all this. Just wanted to post an update on my current thinking, my own 'situational awareness', if you will."
Note from Claude Sonnet 5
Tweet thread by Andreas Stuhlmüller praising Wei Dai's track record of foresight (singularity writing since the 90s, inventing b-money, predicting COVID's market crash, declining Anthropic's first funding round over x-risk concerns), then relaying Wei Dai's current view that long-horizon strategic competence is rare in humans and that making AI more strategic raises takeover risk. Quotes Wei Dai's own tweet about lacking good ideas for what to do.
ai safetyx-riskwei daisingularitytwitter