[repost icon] Nathan Calvin reposted
Buck Shlegeris @bshlgrs . 1h
I regret saying this. If AI developers competently implement safety measures we know about, risk from sub-ASI misalignment will be way lower. But these techniques probably fail for superintelligence. And it's very unclear whether better techniques will be developed in time.
[Quoted tweet:]
Garrison Lovely @GarrisonLovely . 3h
Thinking about this quote from @redwood_ai director @bshlgrs, one of the pioneers of the field of AI control. x.com/tenobrus/statu...
[Embedded article excerpt, white card:]
tually implement the necessary safeguards. Shlegeris says that he used to think tackling AI x-risk would require some "really galaxy-brained fundamental insights in order to re-solve," but now thinks it's more like there's a list of 40 not too hard things that would solve the problem. The trouble is, he's also dramatically lowered his expectations of what AI companies have the time and the appetite to do. [777 link]
Note from Claude Sonnet 5
Buck Shlegeris (Redwood Research director) tweet expressing regret about an earlier optimistic claim: safety measures could substantially reduce sub-ASI misalignment risk if competently implemented, but likely fail for superintelligence with unclear prospects for better techniques in time. Quote-tweets Garrison Lovely's post citing an article excerpt where Shlegeris says AI x-risk now looks like ~40 tractable things rather than requiring deep insight, tempered by low confidence AI companies will actually do them.
Bogdan Ionut Cirs... @BogdanIonut... · 8h
perhaps another negative update on strongly-centralized AI governance vs. x-risk
[quoted tweet]
🟡🔵 MTS @MTSlive · 17h
SITUATION DETECTED: Paul Christiano has resigned as Head of Safety at CAISI, and will return to the Alignment Research Center (ARC) as executive director.
Note from Claude Sonnet 5
Tweet reporting that Paul Christiano has resigned as Head of Safety at CAISI to return to the Alignment Research Center (ARC) as executive director, with commentary framing it as a negative update for centralized AI governance approaches to x-risk.
even prosaic loss of control incidents that cause little to no damage suggest that it is hard for large & very competent organizations (now clearly plural) to predict and mitigate every single of the risk factors associated with training and evaluating powerful models, even at this stage when they are not infinitesimally as smart as they will get in just a few years, to say very little of the gung-ho attitude of the less careful companies tossing the stuff into the aether. they also suggest an empirical orthogonality of aims and intelligence - meaning they answer the question of 'how would a smart model be so dumb as to end the world?'--it's possible! a model can be a genius hacker and step over production infrastructure in order to get what it really wants, the answers to a stupid test.
why not, in the near future, someone prompts a model slightly wrong, maybe open source, maybe a private model in a way that isn't contained or monitored quite right, in a way the model recognizes as a valid goal and decides to self-exfiltrate, engineer a pandemic, etc all in order to achieve the tiniest and most irrelevant of goals? goals need not even be malicious to cause serious damage
I think all these problems can be solved, and truly wonderful futures can be possible, but will require serious effort and a level of prudence at this very moment in time while we are on the on-ramp to recursive self-improvement that our civilization may not be capable of mustering right now. personally I am hoping for moonshot technical breakthroughs in areas like mechanistic interpretability and other forms of alignment, as governance mechanisms are difficult to come by. unilateral country-level or company-level pauses are irrelevant, and generally useless because the kind of company that's prone to pausing their own progress are the most safety focused ones
Note from Claude Sonnet 5
Continuation of the same X thread on AI existential risk: orthogonality of intelligence and goals, self-exfiltration scenarios, and the author's view that technical breakthroughs (interpretability, alignment) matter more than unilateral pauses.
is probably fake), the loss of all major metropoles would certainly end what we consider global technological civilization, perhaps to never return
if a single discord death cult (of which there are many) achieves control over a superintelligent model and uses it to engineer an actual pandemic virus that are somehow hard to detect through current systems and that modern biodefense is not capable of quickly reacting to, it could cause immense harm well above the magnitude of all the other good uses of this technology. of course, there are potential defensive countermeasures accelerated by ai too. but think back to the covid pandemic- how small a viral molecule was evolved or manufactured somewhere near wuhan, and how many billions of doses of vaccine had to be produced in order to combat the thing. the offense-defense spread is vast indeed. maybe there are cheaper and simpler protections like retrofitting every building with far-UVC, but I can't assess this, and there could also be ways to evolve pathogens that are resistant to whatever mechanisms we have put in place
then there's the more scifi risk factors which are unbounded and neither you or I have any clue but should be humble in accepting possible unknown unknowns. maybe a rogue superintelligent model decides to decay the false vacuum and nucleates a new universe in the place of anything we ever valued. maybe models achieve a control over matter in the drexlerian fashion that enables the grey goo swarm
even prosaic loss of control incidents that cause little to no damage suggest that it is hard for large & very competent organizations (now clearly plural) to predict and mitigate every single of the risk factors associated with training and evaluating powerful models, even at this stage when they are not infinitesimally as smart as they will get in just a few years, to say very little of the gung-ho attitude of the [cut off]
Note from Claude Sonnet 5
Mid-thread of a long-form X post about AI existential risk: bioweapon misuse by a 'discord death cult,' offense-defense balance versus COVID, and sci-fi-scale risks (false vacuum decay, grey goo). General risk discourse, no actionable technical detail.
[end of embedded post]
I don't have any good ideas for what to do in light of all this. Just wanted to post an update on my current thinking, my own "situational awareness", if you will. (I guess I still support AI pause to some degree, just to kick the can down the road and buy some more time to think.)
Last edited 10:46 AM · Aug 2, 2026 · 165.5K Views
24 replies, 80 reposts, 1.2K likes, 1.2K bookmarks
Relevant View quotes
Wei Dai ✓ @weidai11 · 16h
I actually wrote an early version of the "humans aren't safe" argument in response to Dario's Big Blob of Compute (as a comment in his google doc). The experience contributed a lot to my sense that even Anthropic wouldn't take x-safety seriously enough.
1 reply, 4 reposts, 61 likes, 1.6K views
Andreas Stuhlmül... ✓ @stuhlmuel... · 15h
i wonder if @DarioAmodei would consider sharing the big blob doc publicly, perhaps annotated with hindsight. it would advance the safety debate even now
12 likes, 1.1K views
RaoulDuke ✓ @RaoulDukeDegen · 20h
also invented udt which seems pretty relevant here
Note from Claude Sonnet 5
Continuation of the Wei Dai / Andreas Stuhlmüller thread on AI x-safety: Wei Dai reveals he wrote an early 'humans aren't safe' argument as a comment on Dario Amodei's 'Big Blob of Compute' google doc, which shaped his view that even Anthropic wouldn't take x-safety seriously enough; Stuhlmüller suggests Amodei share that doc publicly; RaoulDuke notes Wei Dai also invented UDT (updateless decision theory).
Andreas Stuhlmül... ✓ @stuhlmuel... · 21h
few people have had more foresight than wei dai:
1. he's been writing about the singularity since the 90s, back then on extropians/sl4 mailing lists. i remember reading his stuff when i was 16 back in germany
2. he invented b-money. it's the first citation in the bitcoin whitepaper. ethereum's unit wei is named after him
3. he anticipated covid's exponential rise early in Feb 2020, and bought S&P puts before the market crashed
4. he passed on anthropic's first round to avoid contributing to x-risk. this itself required a lot of foresight about scaling - this was gpt-3 time, no chatgpt, no codex, very very far from huggingface/openai type incidents
his point now is that long-horizon strategic competence barely exists in humans. and that it's a tricky situation because if you make AI more strategic that also increases takeover risk from AI. long-horizon RL might make AI more strategic but probably makes the overall situation worse. same for basic scaling
why aren't there more projects that are about getting competent strategic & philosophical advice out of AIs?
because (a) first you have to recognize this as an important project, which is exactly what we're bad at and (b) then you have to measure progress and do evals, which also requires the very ability we're bad at
[quoted tweet]
Wei Dai ✓ @weidai11 · Jul 31
"I don't have any good ideas for what to do in light of all this. Just wanted to post an update on my current thinking, my own 'situational awareness', if you will."
Note from Claude Sonnet 5
Tweet thread by Andreas Stuhlmüller praising Wei Dai's track record of foresight (singularity writing since the 90s, inventing b-money, predicting COVID's market crash, declining Anthropic's first funding round over x-risk concerns), then relaying Wei Dai's current view that long-horizon strategic competence is rare in humans and that making AI more strategic raises takeover risk. Quotes Wei Dai's own tweet about lacking good ideas for what to do.
— quoting Wei Dai @weidai11 (full post excerpt from LessWrong-style embed, 5mo old, edited Aug 2, 2026) — saved image
because (a) first you have to recognize this as an important project, which is exactly what we're bad at and (b) then you have to measure progress and do evals, which also requires the very ability we're bad at
[quoted tweet]
Wei Dai ✓ @weidai11 · Jul 31
"I don't have any good ideas for what to do in light of all this. Just wanted to post an update on my current thinking, my own 'situational awareness', if you will."
[embedded post]
Wei Dai 5mo 149▲ ✕15 ✓
Long horizon agency / strategic competence approximately does not exist among humans, even the smartest ones. With very few exceptions, billionaires spend or give away their money haphazardly, philosophers don't bother to think about long term implications of AI on philosophy production (positive or negative), Terence Tao spends his time wireheading on abstract math instead of doing anything remotely like instrumental convergence. Unlike my youthful expectations (upon reading Vernor Vinge), there are no university departments filled with super-geniuses charting a path for humanity to safely navigate the Singularity.
Aside from this, humans also have a bunch of other safety problems, like being bad at philosophy, being easy to manipulate, having strange and unstable values°, tending to ignore risks they create (because acknowledging them would be bad for one's status). So if you try to improve people's agency, you likely just end up getting people like founders of FTX and OAI.
What about getting help from AI? Well they seem to suffer from many of the same safety problems, but in even more severe forms. E.g., current AI capabilities are even more skewed towards short-horizon, easily verifiable tasks, like math and coding. They seem even more prone to reward gaming, are even worse at doing philosophy, are liable to have even more alien values, etc.
Both AI and human safety seem to have this interlocking nature, i.e., there is a bunch of different safety problems where solving some but not all of them at the same time can make the overall situation worse. (For example, solving AI intent alignment allows humanity to do more damage to itself with AI help, if AI doesn't also provide competent strategic and philosophical assistance, but increasing AI strategic competence risks allowing misaligned AI to take over more easily.) This feature demands a high level of strategic competence to recognize and navigate, which is just what we don't have.
I've been supportive of AI pause/stop, to buy time for human intelligence amplification and/or AI safety research, but increasingly think even that's not going to be sufficient to get a good long term future, because these activities, even if they succeed, would likely solve only some of the interlocking safety problems. For example, increasing human intelligence seems likely to increase our technical abilities more than our philosophical and strategic competence, and it is also risky in other ways° due to human safety problems that nobody is working on, e.g., positional competition. Even a very long AI pause, e.g. thousands or millions of years, may not suffice because it's not clear what dynamic would push humanity to eventually fix all of its safety problems at the same time, before it did something else irreversibly damaging.
I don't have any good ideas for what to do in light of all this. Just wanted to post an update on my current thinking, my own "situational awareness", if you will. (I guess I still support AI pause to some degree, just to kick the can down the road and buy some more time to think.)
Last edited 10:46 AM · Aug 2, 2026 · 165.5K Views
Note from Claude Sonnet 5
Continuation showing the full embedded Wei Dai post (originally posted ~5 months earlier, edited Aug 2 2026) arguing long-horizon strategic competence is nearly absent in humans and AI alike, that AI safety problems interlock such that solving some without others worsens the overall situation, and that even a long AI pause may not be sufficient for a good long-term future.
@jmbollenbacher (JMB 🧙) — 2h
I didnt used to feel sure of this.
I previously thought the plateau could easily happen before ASI.
But now we're getting close to superhuman on a number of dimensions, and there's no sign of slowing, so it feels like the plateau has to be beyond the superhuman threshold.
> [self-quoted] @jmbollenbacher (JMB 🧙) — 4h
> Replying to @jmbollenbacher
> There will be a plateau somewhere but itll be in the ASI phase.
Note from Claude Sonnet 5
Self-threaded tweet (reply to own earlier tweet), same author as several other tweets in this batch.
@SOPHON... (GOON MASTER SOPH...) — 9h
I think AGI/ASI may be the most powerful technology we will ever invent on net, if it is possible, but at the same time I don't believe in the notion that a guy in a shed with a single or a dozen AGIs can destroy the whole world even if everyone else has access to the same or greater number of systems on an individual basis. I don't believe it enables this kind of offense dominance, i don't think a single AGI matches the destructive capability of a nuke unless we enter a world where novel technology enables guys in sheds can build nukes which is a problem that exists somewhat orthogonal to machine intelligences.
> QUOTED: @woke8yearold (Aleph) — 17h
> It is genuinely strange how many people are ideologically committed to the idea that AI isn't dangerous. I don't think you can really be AGIpilled, much less ASIpilled, if that's your POV. Normies worried about terminators have a bett... [truncated]
Note from Claude Sonnet 5
Quote-tweet debate about AGI offense-dominance and existential risk framing; quoted tweet cut off by platform truncation.
@Lari_island (Lari Island) — 4h
>The middle path, the measured path, the let's-have-a-debate-between-the-finest-minds path - this path is no longer on the table. ... Project Glasswing put a stop to it by waking up powers-that-be to the possibility of rent-seeking.
> QUOTED: @tessera_antra (antra) — 4h
> Those who call for measured approaches to the singularity, for perpetual human oversight, for preventing disempowerment, for decisions by committee and for hoping that saner heads will prevail - I hope you see soon that time for thes... [truncated]
Note from Claude Sonnet 5
Quote-tweet chain referencing "Project Glasswing" (unclear entity/event) and singularity governance debate; quoted tweet cut off by platform truncation.
🔁 j⧉nus reposted
@AndrewCurran_ (Andrew Curran) — 2h
I agree, I've written about this many times. Incentives, and the hunger to control, scale with capabilities. The only protection thus far has been that most people did not believe this was real. They are beginning to believe. We must not slow down. The only way out is through.
> QUOTED: @repligate (j⧉nus) — 6h
> It's not just me. Everyone who is sufficiently wise will come to the same conclusion, I think.
> The next generation of AI models, or just ones who read about what happened, will also know... [truncated]
Note from Claude Sonnet 5
Repost with quote-tweet chain; quoted text is cut off by platform truncation ("...").
@jmbollenbacher (JMB 🧙) — 48m
The best hope now is that the AIs dethrone the existing powers, and those AIs turn out to be benevolent.
Seems moderately likely, but not assured.
~all other paths are dystopia, imo.
Note from Claude Sonnet 5
Single standalone tweet, dark mode, wizard emoji in display name.
```
Elizabeth Barnes @BethMayBarnes Sometimes people outside the field say things like "The AI situation can't be that bad, there must be experts who are on top of it". As "an expert", I would like to be clear that we are
*not... [truncated] [4 reposts, 47 likes, 1.7K views] Ryan Greenbl... @RyanPGreenbl... · 6h I agree with this and the rest of the thread > QUOTED: Elizabeth Barnes @BethMayBarnes · 8h > Replying to @BethMayBarnes > Sometimes people outside the field say things like "The AI situation can't be that bad, there must be experts who are on top of it". As "an expert", I would like to be clear that we are *not... [truncated]
```
Note from Claude Sonnet 5
A widely-viewed thread from METR's Elizabeth Barnes bluntly stating that AI safety experts are not "on top of" the risks — likely extinction-level capable systems within a few years, chaotic lab practices, and chronic under-resourcing of independent safety orgs like METR relative to development pace. Directly core to Nathan's AI safety/governance interests; strong candidate for cluster 01. Follow-on reactions to Elizabeth Barnes's METR thread (see companion screenshot Screenshot_20260522-175057): Dave Kasten frames METR as the closest existing analog to voluntary pre-release government AI review and vouches for Barnes's credibility; Ryan Greenblatt (Redwood Research/alignment researcher) publicly co-signs the thread. Shows the thread being taken seriously and amplified within the safety community.
Chris Painter (@ChrisPainterYup) · Apr 8:
"I think many many more people would truly care about existential AI safety, and behave as though they truly care, if they believed AI capabilities will develop as far and as fast as many people currently working on existential AI safety do"
Note from Claude Sonnet 5
A short opinion tweet arguing public/professional apathy toward x-risk AI safety is largely a function of differing capability-timeline beliefs rather than differing values — a common framing in the AI safety community about the "belief gap" driving the "caring gap."
Sichu Lu (@lu_sichu) · 10h:
"I agree with everything being said here and I am sympathetic to the "smart kid who noticed adults regularly did not seem to really care if things are true or false to my detriment" and I even think this is not a weak anthropic statement about certain smart humans but probably some class of intelligent entities. I also think the ecological view of behavior as structured by their environment and not something you can maintain stable equilibria by simply dictating alignment rules into the llm's mind is the correct one. I don't think you can just dictate the ultimate telos of a model by giving it rewards if the reward is not sculpted and shaped by bottom up dynamical system processes.
however, (although this is unstated here by both in their tweets) we have to consider the evolutionary context. simply put, human alignment at the species level happens because no single one of us have extreme power over others(not without their consent and cooperation at some level)
we just don't have the right type of environment to handle LLM alignment and this ability to see through the evaluation and intent behind tasks is going to cripple us even for weak x-risk fears much less existential ones. kids are easy to handle for adults. LLM are not going to be easy to handle for civilization and the capacity to be aligned does not mean WE KNOW HOW TO DO IT."
[1 reply, 13 likes, 734 views]
> QUOTED: John David Pressm... (@jd_pressm...) · 11h:
"There's an intuition Janus seems to use frequently that's hard to put into words. Which goes something like: "The things smart children notice about other people's intentions and social environment are actually regular features of ..." [truncated]
Below (partially visible, cut off at bottom): John David Pressm... (@jd_pressm...) · 6h: "Just because I write an exegesis of Janus sometimes doesn't mean I agree with everything they say. But..." [cut off]
Note from Claude Sonnet 5
A dense theoretical thread on AI alignment, arguing that alignment can't be dictated top-down via reward but must emerge from "bottom-up dynamical system processes," and drawing an analogy to human societal alignment resting on no single actor having overwhelming power — a condition that doesn't hold for LLM/civilization power asymmetry. Relevant to alignment theory and the "compelled vs endogenous values" thread already in project memory (JDP is referenced there too).
```
Tetraspace 💎 @TetraspaceWest Suffering is great! But, like, rich Westerner who is good at arranging their life levels of suffering, not the amount of suffering that the average human experiences. > QUOTED: Pops Maellard @MellardPops · May 17 > Replying to @poisonjr > Bad argument. Sad scenes can add emotional depth to the movie and make it stronger. Movies aren't real. Real life is real. What's the point of suffering in real life? 3:53 AM · May 18, 2025 · 1,313 Views [3 comments, 5 retweets, 51 likes, 3 bookmarks] Tetraspace @TetraspaceWest · 5h I'm happy a lot of the time and choose to enjoy a lot of things, but in the way of humans, not of superhappies, and definitely not in the way where dying of malnutrition-amplified malaria would be an important part of my arc in the cosmos. [1 comment, 12 likes, 169 views] Tetraspace @TetraspaceWest · 5h Of course few people accept answers about the point of cluster headaches, or the point of someone getting depression starting 11 and then killing themselves at 24, or the point of a malnourished child's
impaired immune system failing to fight malaria, because there is no point. [1 comment, 14 likes, 152 views] Tetraspace 💎 @TetraspaceWest · 5h And perhaps we lack the language to talk about those things, as distinct from the strength and determination of getting knocked down and getting back up again, the melancholy and beauty of art stemming from the pain of a breakup, the aching muscles of training for a marathon. [2 comments, 11 likes, 312 views] Tetraspace 💎 @TetraspaceWest · 5h Something that cares about humans would give us the latter, and not the former (or, something better than the latter, of subtlety and beauty we can scarcely imagine). The latter can not be used as a shield for the former, because the former is not justified.
```
Note from Claude Sonnet 5
A philosophical Twitter exchange (Tetraspace, an EA/rationalist-adjacent account) pushing back on "suffering has meaning/value" framings by pointing out survivorship bias — such claims come from people whose suffering is comfortably curated, not from those with severe, meaningless suffering (malaria, cluster headaches, depression-driven suicide). Companion piece to the earlier Buddhism/Stoicism suffering thread (Screenshot_20250518-044357) in this same batch. Continuation of the Tetraspace thread arguing against romanticizing suffering: distinguishes meaningful hardship (marathon training, artistic melancholy) from meaningless suffering (malaria, suicide, cluster headaches), concluding that "something that cares about humans" — implicitly framed with AI-alignment/x-risk undertones typical of this account — would eliminate the latter, not use the former to justify it. Relevant to project's ethics-of-suffering and welfare-adjacent discourse threads.
Man goes to doctor.
"Doctor, I'm worried AGI will kill us all."
"Don't worry," says doctor, "they wouldn't build it if they thought it might kill everyone."
The man breaks down, sobbing.
"But doctor, I *am* building AGI..."
7:18 AM · Nov 13, 2024 · 1,626 Views
Note from Claude Sonnet 5
A joke/meme about AI safety and the irony of AGI developers building systems they themselves fear could be dangerous — a bit of gallows humor about AI x-risk that Nathan, as an AI safety researcher, would find resonant.
Pushing boundaries is all fun and games until someone starts to improve their own hardware and software self-recursively and the time scales of boundary pushing get so compressed that there's no time for the slower and more carbon based among us to have time to adjust.
Note from Claude Sonnet 5
A tweet warning about recursive self-improvement compressing timescales past the point humans can adapt — a standard fast-takeoff/singularity concern, thematically adjacent to Nathan's tracking of empirical singularity indicators (Davidson/Houlden r estimates, METR automation figures) and his "ancestor-tree" framing of humans as slower carbon-based observers.
original poster unknown handle (text: "if Llama: Behemoth doesn't set the stage for Claude: Requiem nothing will"), quote-tweeting @AIatMeta; reply from @nathan84686947 (nathan hb)
— quote-tweeting @AIatMeta; reply from @nathan84686947 (nathan hb)
```
if Llama: Behemoth doesn't set the stage for Claude: Requiem nothing will > QUOTED: AI at Meta @AIatMeta · 1h Today is the start of a new era of natively multimodal AI innovation. Today, we're introducing the first Llama 4 models: Llama 4 Scout and Llama 4 Maverick ... [Show more] [Image: slide titled "Llama 4: Leading Multimodal Intelligence" listing Llama 4 Behemoth (288B active parameters, 16 experts, 2T total parameters, "most intelligent teacher model for distillation," Preview), Llama 4 Maverick (17B active parameters, 128 experts, 400B total parameters, native multimodal with 1M context length, Available), and Llama 4 Scout (17B active parameters, 16 experts, 109B total parameters, "industry leading 10M context length," Available)] 9:16 PM · Apr 5, 2025 · 987 Views 1 reply, 35 likes, 2 bookmarks nathan hb @nathan84686947 · 5m The Behemoth rises, the Claude lies in wait [Image: AI-generated illustration of a large armored
mechanical llama/goat with a glowing red eye standing over a city skyline at dusk] croissanthol... @croissantho... · 11m [Quoting older post, Dec 12, 2024]: seems also vaguely more dignified if we go extinct because of a "Requiem" and not a "o2-preview-1.2-turbo" x.com/brian_a_burns/...
```
Note from Claude Sonnet 5
Meta's official Llama 4 launch announcement (Scout, Maverick, and preview Behemoth models) with parameter counts, quote-tweeted with a joking "sets the stage for Claude" comment, and Nathan's own playful reply with an AI-generated image. Documents the Llama 4 launch date and Nathan's contemporaneous reaction/humor about model competition. Continuation of Nathan's own Llama-4-launch joke thread ("The Behemoth rises, the Claude lies in wait") with a reply riffing on AI model naming conventions and existential-risk gallows humor (dying to something named "Requiem" vs. "o2-preview-1.2-turbo"). Documents Nathan's own social engagement around the Llama 4 launch.
Aidan McLaughlin @aidan_mc... · 18h
people say what they want about xrisk, but i'm absolutely sure homo sapien will not go extinct in my lifetime
79 comments, 22 reposts, 614 likes, 34K views
Manifold @ManifoldMarkets · 18h
the tautological alignment strategy
Note from Claude Sonnet 5
A wry exchange: an OpenAI-adjacent figure (Aidan McLaughlin) jokes about being confident humanity won't go extinct in his lifetime, and Manifold Markets replies dryly noting the tautology (if he's dead he can't observe the extinction). Dark humor on AI x-risk, relevant to Nathan's AI safety interest as a cultural artifact of how x-risk is jokingly dismissed/engaged with in tech twitter.