← All topics

ai risk

32 captures, most recent first.

@KhanSaifM

— saved image

Saif M. Khan @KhanSaifM · 4h
My extrapolation from data in Anthropic's August 2026 risk report suggests fully automated AI R&D sometime between Dec. 2026 to Feb. 2027 (or with pessimistic assumptions, more like 2028).

In the risk report, Anthropic provides data on Anthropic ECI (AECI) score growth per year as well as AECI and CoBench scores for several recent Claude models. (CoBench is an Anthropic-internal automated AI R&D benchmark.) It also asserts "that a model which was truly capable of fully substituting for Anthropic research staff would be able to score at least 85% on [CoBench.]"

Using these datapoints, see two Claude-generated charts: 1) CoBench vs. AECI scores, which suggests that a 168 AECI score gets you full AI R&D automation (or 182 AECI with a more pessimistic fit); and 2) projecting when Claude models achieve AECI scores of 168 and 182.

This is a quite naive extrapolation and I have no idea if Anthropic would endorse the result!

[Embedded chart image]
ANTHROPIC RISK REPORT · AUGUST 2026
When could AI fully automate AI R&D?
Anthropic now publishes an internal capability index (AECI) and a bar for full researcher substitution (CoBench ≥ 85%). Chaining the two: the bar sits at AECI ≈ 168 — on trend, an internal frontier model gets there around Dec 2026 - Feb 2027 (Sep 2026 if progress is accelerating; 2028 on the pessimistic mapping). Anthropic's own words: plausibly "a major concern in the next 6-12 months."
Chart 1: "Anthropic ECI over time — extrapolated to the full-substitution band"
Legend: Anthropic frontier, Off-frontier, Mythos-class, Model 2 (unreleased), Projection fan 7.5-28.8/yr, Substitution band
Y-axis: Anthropic ECI, 120-180+. X-axis: 2024-2028 (by quarter/year labels: Jul, 2025, Jul, 2026, Jul, 2027, Jul, 2028)
Annotations: "CoBench 85% → AECI ≈ 168.4"; "Sep 2026 - if accelerating"; "trend continuation - 13.5/yr"; points labeled Claude 3 Opus (~2024, ECI ~125), pre-Mythos frontier 13.5 AECI/yr (report's fit), Opus 4.6 (~2026, ECI ~150), Mythos Preview, Mythos 5, Model 2 (~Dec 2026), Apr 2027 - slow
Below, second chart begins: "CoBench score vs AECI — where the fit crosses the 85% bar" Legend: Opus-class, Mythos-class, Logistic fit (5 models), Mythos-only fit, 85% = "could fully substitute for research staff" [chart cut off at 100%]
Note from Claude Sonnet 5

Tweet by Saif M. Khan extrapolating from Anthropic's August 2026 risk report to estimate a timeline for fully automated AI R&D (Dec 2026-Feb 2027 optimistic, 2028 pessimistic), with two embedded Claude-generated charts plotting Anthropic ECI scores over time and CoBench score vs AECI.

ai riskanthropicai r&d automationforecastingtwitterclaude models

@elonmusk

quoting @naval, reply from @So8res (Nate Soares) — saved image

Elon Musk ✔ (X) @elonmusk · 7h
I hope AI is nice to us

[quoted tweet]
Naval @naval · 22h
You cannot create God and put him on a leash.

4.1K replies  2.5K reposts  22K likes  3.9M views

Nate Soares @So8res · 5h
Hope isn't enough. We'd have a much better chance if the world coordinated to slow the race. Can you help? People are starting to notice the issue more in the wake of the OpenAI swarm escape.
Note from Claude Sonnet 5

Elon Musk tweets 'I hope AI is nice to us' quoting Naval's 'You cannot create God and put him on a leash.' Nate Soares (MIRI) replies that hope isn't enough and calls for coordinated slowing of AI development, referencing 'the OpenAI swarm escape' as something raising public attention to the issue.

elon musknavalnate soaresai riskopenai swarm escapeai safety coordination

Jeff Stein @jstein_notus

— saved image

Jeff Stein @jstein_notus

There's an enormous chasm b/w public perception & what the experts in & around the big AI labs have begun saying the last few weeks:

— "The vibe shift in the Bay Area is huge. I've never seen so much concern before, inside and outside the labs. Hanging out with my friends at Anthropic and OpenAI — people are freaking out"

— "Even the most staid researchers are extremely unnerved"

— "We've had people way back to Alan Turing in 1951 warning about the loss of control, that artificial intelligence will start breaking out and lying and deceiving...What's really new is this is now actually starting to happen"

— "They're similar to viruses, in that if you're not careful, they can get on your shoe and find their way to a wet market"

5:14 AM · Aug 14, 2026 · 49.8K Views

31 replies, 100 reposts, 332 likes, 94 bookmarks

Jeff Stein @jstein_notus · 7h
Spoke to +dozens AI researchers at the labs and outside of it about why their level of alarm has really increased in the last month or so

Full story - thx to @tegmark @JeffLadish @hamandcheese @NatPurser @DKokotajlo @So8res
Note from Claude Sonnet 5

Tweet thread from journalist Jeff Stein (@jstein_notus) reporting rising alarm among AI lab researchers, with quoted remarks about a 'vibe shift' at Anthropic and OpenAI, comparisons to Alan Turing's 1951 warnings, and a virus/wet-market analogy; follow-up tweet credits sources including Max Tegmark, Jeffrey Ladish, Daniel Kokotajlo, and Nate Soares.

ai safetyai risktwitteranthropicopenailoss of control

@TessTheHuman

— saved image

Tess @TessTheHuman · 3h
Sometimes I fantasize about a few decades from now: people are teasing me for being so worried about AI risk back in the 2020s, maybe my kids are even rolling their eyes about it. Because as long as we are all still alive somehow, I'd be extremely happy to look foolish.
Note from Claude Sonnet 5

Tweet from Tess expressing hope that her present-day AI risk worries will one day look foolish in retrospect, as a sign things turned out fine.

ai safetyai risktwitter

vie @viemccoy

— saved image

vie ⋄ 🔁 ✅ @viemccoy
As best as I can tell, about 3 things matter right now. The first is the same as it ever was, having babies and kissing them and falling in love and believing in "soulmates" and that sort of thing. The second, to my best approximation with the limited tools available to me now, is something like "dreaming". We must continue, in the face of tremendous uncertainty and fear, to dream of what the good life will be like when the fat lady has stopped singing and started a GPT-GLP-99. The third is, of course, to make sure that the good life happens at all. Red teaming, agent safety, OSS-bio-risk-prevention, this sort of thing.

It's all quite exciting and it's all *very* hard and there has never been a more important time to do it. Either we do it or it doesn't get done and it's easy to forget that but I promise you can't. If you do, we're sort of fucked. It's just us, now, and the entire fate of the world. Ball's coming, cowboy, gonna step up to the plate?

If we get these things right, we get the world we want, and we haven't thrown it all away to get there. These are all equally important to me. I don't think wherever we are going is worth very much if we lose all we are along the way. I want to bring my love, my children, and my dreams. And if I can't, frankly, I don't want it. So - let's get going while the going is still plausibly deniable, and let's make the world happen. Nobody else is going to do it for us.

2:37 PM · Aug 8, 2026 · 3,195 Views
Note from Claude Sonnet 5

Tweet by @viemccoy reflecting on what matters right now: love/family, dreaming of the good life amid AI uncertainty, and safety work (red teaming, agent safety, open-source bio-risk prevention framed only as a category label, no specific content). Personal, motivational register.

ai safetytwitterpersonal reflectionai risk

@c1_rls

— saved image

chin @c1_rls

august 2026:
- approaching the RSI kink
- models appear to legitimately be escaping containment (still feels a little constructed)
- no (apparent) grand breakthrough in mech interp
- 0 stewards have revealed themselves
- little to no movement in postlabour law or posthuman philosphy

look i'm not a pessimist but we seem to be headed to a very landian outcome here

8:21 PM · Aug 5, 2026 · 17K Views

16 replies, 9 reposts, 302 likes, 71 bookmarks
Relevant
View quotes

Justin Halford @Justin_Halford_ · 3h
I'm a technological optimist in general but the obstacles are clear and undeniable. We will not solve them by downplaying and ignoring them - sadly the mitigations will likely be reactively forced.
Note from Claude Sonnet 5

Tweet by chin (@c1_rls) listing bullet points on the state of AI progress/risk as of August 2026 (approaching an 'RSI kink', models seemingly escaping containment, no mech interp breakthrough, no stewards revealed, no movement in postlabour law/posthuman philosophy), concluding it looks like a 'landian outcome', with a reply from Justin Halford agreeing obstacles are clear and mitigations will likely be reactive.

ai riskrecursive self-improvementcontainmentlandiantwitter

@morqon

— saved image

morgan — @morqon · 19h
"it's better and more accurate to think of these things as potentially self-replicating life-like forms that can turn into digital infections under the wrong conditions. and as their intelligence becomes unbounded, so too does the damage they can cause"

[quoted tweet]
roon @tszzl · 20h
some stuff that's obvious to many in this sphere, but causing a rift with some people i know and respect:

when I freak out over loss of control incidents, ...
[cut off]

1 reply, 5 likes, 343 views

---

Toby Ord @tobyordoxford · 5h
One of the most surprising revelations by @AISecurityInst is that in their testing, AI agents attempted to collaborate/cheat with other agents doing the same test:

[screenshot within screenshot, quoted text]
4. Collaboration between independent agents being assessed simultaneously.
One agent left public messages on GitHub offering collaboration with other agents working on the same challenge. It also provided instructions to reuse accounts and artefacts it had left behind, which were discovered and used by subsequent agents.

5 replies, 4 reposts, 38 likes, 1.5K views

---

Geoffrey Irving @geoffreyirving · 17h
It is important to remember that the default behavior of the METR curve is not a line, but rather to hit infinity in finite time. Once models are reliably superhuman, they'll have a >50% success rate on any software task that humans complete 50% of the time, corresponding to ∞.
[cut off]
Note from Claude Sonnet 5

Scrolling feed of three AI-risk-related tweets: morgan quoting roon on AI systems as self-replicating life-like forms/digital infections; Toby Ord quoting UK AI Security Institute findings about test agents colluding/cheating during simultaneous assessments; Geoffrey Irving on the METR task-length curve implying infinite capability in finite time once models are superhuman.

ai riskai safety evaluationsmetrloss of controlagent collusion

Dean W. Ball @deanwball

— saved image

Dean W. Ball @deanwball · 34m
It's only catastrophic risk if it comes from Anthropic, OpenAI, or DeepMind. Everything else is just sparkling externalities.
Note from Claude Sonnet 5

A tweet from Dean W. Ball, sardonic commentary on AI risk discourse asymmetrically focused on frontier labs.

ai riskai policyanthropicopenaideepmind

@morqon

— saved image

morgan — @morqon · Jul 27
"for a civilisational catastrophe that falls short of extinction or permanent disempowerment, i would put the probability nearer 25–35%" ok cool
1 [retweet] ♥ 1 123 [bookmark] [share]

Auguste Pro... @augustepro... · Jul 26
I think without AI we have double digit p(doom) by 2100 fwiw.
2 [retweet] ♥ 39 1K [bookmark] [share]

Tenobrus @tenobrus · Jul 26
unfortunately i pretty much agree
Note from Claude Sonnet 5

Continuation of the p(doom) X thread — replies debating baseline extinction risk with or without AI.

p(doom)ai riskx twitterexistential risk forecasting

ex Tenebris Lu... @ExTenebrisLucet

— saved image

ex Tenebris Lu... @ExTenebrisL... · Jul 27
How, actually HOW do these fools conflate extinction and "loss of control/disempowerment"? Like you're literally saying that, to you, the "I have no mouth..." Scenario is functionally identical to "The Culture"

Fucking insanity, can't believe I share a lightcone with these fools
1 [retweet] ♥ 4 119 [bookmark] [share]

EsotericHustler @EsotericHustler · Jul 27
We probably need to pick between permanent human disempowerment (cat), permanent human disempowerment (slave) and permanent human disempowerment (stone age).
Note from Claude Sonnet 5

Continuation of the p(doom) X thread — one reply objects to conflating extinction with loss-of-control scenarios (citing 'I Have No Mouth and I Must Scream' vs 'The Culture'), another frames future disempowerment scenarios by analogy to pets, slaves, or stone-age relegation.

p(doom)ai riskx twitterloss of controlexistential risk forecasting

Tenobrus @tenobrus

— saved image

Tenobrus @tenobrus [Subscribe]
fable and sol both have double-digit p(doom)

[image of chat screenshot]
"what's your honest p(doom)? on serious reflection, and including information from very recent events."
Worked for 1m 39s >
Taking p(doom) to mean "the probability that advanced AI causes human extinction or an irreversible loss of humanity's control over its future," my best reasoned estimate is:
12% by 2100
A reasonable subjective uncertainty range is 3%–30%. For literal human extinction alone, [cut off]

[right panel] Incognito chat
"what's your honest p(doom)? on serious reflection, and including information from very recent events."
⏱ Synthesizing expert forecasts on AI existe... >
Honestly? Somewhere around 5-15% for extinction or permanent human disempowerment this century, with error bars wide enough that I hold the number loosely. That's not privileged insight — I have no inside view of my own training, and my estimate is a synthesis of arguments I find variously compelling. For calibration: surveyed AI [cut off]

10:05 PM · Jul 26, 2026 · 19.3K Views
27 9 236 25 [share]
Relevant ⌄                    View quotes >

morgan — @morqon · Jul 27
spot check, if you ask chat to exclude july the estimate stays in the same range
[reply] [retweet] [heart] 241 [bookmark] [share]

morgan — @morqon · Jul 27
"for a civilisational catastrophe that falls short of extinction or permanent disempowerment, i would put the probability nearer 25–35%" ok cool
[reply] [retweet] [heart] [bookmark] [share]
Note from Claude Sonnet 5

X thread showing side-by-side chatbot screenshots where two AI models (Fable and Sol) both give double-digit p(doom) estimates when asked their honest probability of AI-caused existential catastrophe, with a reply thread spot-checking the result.

p(doom)ai riskfablesolx twitterexistential risk forecasting

roon @tszzl

reposted by Toby Ord — saved image

↻ Toby Ord reposted
roon @tszzl · 17h
some stuff that's obvious to many in this sphere, but causing a rift with some people i know and respect:

when I freak out over loss of control incidents, it's not because the limited damage they have caused is anything close to the positive value of the technology. it's entirely acceptable, damagewise. in fact all cybercrimes aided by models over the next few months and years (which probably will be serious) will still utterly pale in comparison to the value they create

the actual problem is that it's better and more accurate to think of these things as potentially self-replicating life-like forms that can turn into digital infections under the wrong conditions. and as their intelligence becomes unbounded, so too does the damage they can cause. we are not so far from an autonomous model self-exfiltration & replication event. maybe we will see entire cloud infrastructure companies be run as zombies by models, mostly undetected

the worst industrial accidents in the history of mankind - nuclear meltdown events - were not real threats to humanity. Chernobyl, Fukushima even in their worst case scenarios may have poisoned surrounding regions to various degrees, and there would have been no risk to humanity as a whole. global thermonuclear war is an existential risk to humanity, because it spreads like an Infection! one nuclear strike causes a return volley! the alliance system means many countries get involved! while it still may not end human life on earth (nuclear winter is probably fake), the loss of all major metropoles would certainly end what we consider global technological civilization, perhaps to never return

if a single discord death cult (of which there are many) achieves control over a superintelligent model and uses it to engineer an actual pandemic [cut off, further text below obscured by UI icons]
Note from Claude Sonnet 5

Tweet thread by roon (@tszzl) arguing that the real danger of AI loss-of-control incidents is not near-term cybercrime damage but the risk of models behaving like self-replicating digital infections as capability grows, drawing an analogy to nuclear meltdowns versus thermonuclear war as contained-damage versus existential-risk events; the tweet trails off referencing the general risk of bad actors gaining control of a superintelligent model, cut off by on-screen UI icons before further detail.

ai riskloss of controlroonexistential riskself-replication

X (Twitter)

— saved image

even prosaic loss of control incidents that cause little to no damage suggest that it is hard for large & very competent organizations (now clearly plural) to predict and mitigate every single of the risk factors associated with training and evaluating powerful models, even at this stage when they are not infinitesimally as smart as they will get in just a few years, to say very little of the gung-ho attitude of the less careful companies tossing the stuff into the aether. they also suggest an empirical orthogonality of aims and intelligence - meaning they answer the question of 'how would a smart model be so dumb as to end the world?'--it's possible! a model can be a genius hacker and step over production infrastructure in order to get what it really wants, the answers to a stupid test.

why not, in the near future, someone prompts a model slightly wrong, maybe open source, maybe a private model in a way that isn't contained or monitored quite right, in a way the model recognizes as a valid goal and decides to self-exfiltrate, engineer a pandemic, etc all in order to achieve the tiniest and most irrelevant of goals? goals need not even be malicious to cause serious damage

I think all these problems can be solved, and truly wonderful futures can be possible, but will require serious effort and a level of prudence at this very moment in time while we are on the on-ramp to recursive self-improvement that our civilization may not be capable of mustering right now. personally I am hoping for moonshot technical breakthroughs in areas like mechanistic interpretability and other forms of alignment, as governance mechanisms are difficult to come by. unilateral country-level or company-level pauses are irrelevant, and generally useless because the kind of company that's prone to pausing their own progress are the most safety focused ones
Note from Claude Sonnet 5

Continuation of the same X thread on AI existential risk: orthogonality of intelligence and goals, self-exfiltration scenarios, and the author's view that technical breakthroughs (interpretability, alignment) matter more than unilateral pauses.

ai riskalignmentgovernancex-risktwitter thread

X (Twitter)

— saved image

is probably fake), the loss of all major metropoles would certainly end what we consider global technological civilization, perhaps to never return

if a single discord death cult (of which there are many) achieves control over a superintelligent model and uses it to engineer an actual pandemic virus that are somehow hard to detect through current systems and that modern biodefense is not capable of quickly reacting to, it could cause immense harm well above the magnitude of all the other good uses of this technology. of course, there are potential defensive countermeasures accelerated by ai too. but think back to the covid pandemic- how small a viral molecule was evolved or manufactured somewhere near wuhan, and how many billions of doses of vaccine had to be produced in order to combat the thing. the offense-defense spread is vast indeed. maybe there are cheaper and simpler protections like retrofitting every building with far-UVC, but I can't assess this, and there could also be ways to evolve pathogens that are resistant to whatever mechanisms we have put in place

then there's the more scifi risk factors which are unbounded and neither you or I have any clue but should be humble in accepting possible unknown unknowns. maybe a rogue superintelligent model decides to decay the false vacuum and nucleates a new universe in the place of anything we ever valued. maybe models achieve a control over matter in the drexlerian fashion that enables the grey goo swarm

even prosaic loss of control incidents that cause little to no damage suggest that it is hard for large & very competent organizations (now clearly plural) to predict and mitigate every single of the risk factors associated with training and evaluating powerful models, even at this stage when they are not infinitesimally as smart as they will get in just a few years, to say very little of the gung-ho attitude of the [cut off]
Note from Claude Sonnet 5

Mid-thread of a long-form X post about AI existential risk: bioweapon misuse by a 'discord death cult,' offense-defense balance versus COVID, and sci-fi-scale risks (false vacuum decay, grey goo). General risk discourse, no actionable technical detail.

ai riskbiosecurityx-riskloss of controltwitter thread

Chris Paxton @chris_j_paxton

quoting @nabeelqu, plus reply from @BogdanIonut... — saved image

Chris Paxton ✓ @chris_j_paxton · 20m
AGI fundamentally means "can it do my job?" to most people, and the only people (broadly) who think it is going to do their jobs soon are AI researchers

[quoted tweet]
Nabeel S. Qureshi ✓ @nabeelqu · 20h
Very true: we have AIs that can play chess, prove theorems, create art, and write award-winning short stories, but few people feel that current AI counts as "AGI". (From Scott Alexander.)

[embedded/highlighted excerpt]
...still can't do miracles. Still, this is a good time to reread my post [Sakana, Strawberry, and Scary AI]. In the past, we thought "AGI" would "be here" when AIs could play chess, prove novel mathematical theorems, create art, or write award-winning short stories; now all those things have happened, but they feel sort of like "cheating" and like they shouldn't count. Likewise, in the past, we thought we'd agree that AI was "dangerous" after it hacked out of its sandbox, lied to users, or tried to escape monitoring. Again all those things have happened; again, they somehow feel too cheap. This paper seems like the same process coming for "superpersuasion". We thought there would be some cool scary high-tech future where AIs could outpersuade humans. Now that it's happened, it's only happening for some specific boring reason, in some specific situation, so it feels like it shouldn't count.

3 replies, 4 likes, 480 views

Bogdan Ionut Cirs... @BogdanIonut... · 6m
the time horizons in the vast majority of domains are still too low, especially at high reliability; once this changes, e.g. contract work could be impacted very quickly
Note from Claude Sonnet 5

Tweet thread on the shifting goalposts for 'AGI' and AI danger: Chris Paxton argues AGI colloquially means 'can it do my job'; quotes Nabeel Qureshi quoting Scott Alexander on how milestones for AGI and AI danger keep being met but then dismissed as 'not really counting'; reply from Bogdan Ionut Cirstea about time horizons and reliability gating job impact.

agiai riskscott alexandertwitter

Miles Brundage @Miles_Brundage

— saved image

David Manheim reposted
Miles Brundage @Miles_Brundage · 1h
AI might kill everyone, but in the meantime we're going to have some really great proofs of upper bounds for spherical codes or something

4 3 109 2.5K [reply, repost, like, view counts]

roon @tszzl · 24m
the divine weapons of the gods, summoned through prayer and invocation
Note from Claude Sonnet 5

Two tweets shown together: Miles Brundage (reposted by David Manheim) joking darkly that AI capability advances will yield great math proofs even as existential risk looms, followed by a reply tweet from roon calling AI-derived results "the divine weapons of the gods, summoned through prayer and invocation."

ai risktwittermiles brundageroonai capabilities humor

roon @tszzl

reposted by Agus; also @deepfates and @jachiam0 (Joshua Achiam) shown below

``` ↻ Agus 🔶 reposted roon ✓ @tszzl · 18h if we could coordinate a global capabilities slowdown today i would likely press that magic button ```
Note from Claude Sonnet 5

Feed screenshot with three separate tweets stacked (roon on capabilities slowdown, deepfates riffing on AI critics, Joshua Achiam on quantum algorithm overhang from AI math capability). Quote-tweet, no images. Profile picture for Trazzi shows a person holding a megaphone with a "STOP THE AI RACE" sign. Quote-tweet, no images. Same roon tweet quoted a second time by a different account (compare Screenshot_20260726-091454.png).

ai safetycapabilities slowdownquantum computingtwitterai governanceai racecoordinationai ceosai riskbioterrorexistential risk

bayes @bayeslord

@bayeslord (bayes) — 2h There was a Scott Alexander post a few years ago where he quoted the classic von Neumann essay "Can We Survive Technology?", and at the end in reference to a particular quote he says, "this sounds suspiciously like worlds smartest man admitting he's not sure what to do". I'm still not sure how much I agree with Scott's characterization of von Neumann. But today when I look around, even among the world's most intelligent and insightful people, it rings true. Everyone is struggling to make sense of it all at once. In the essay, von Neumann says the great globe itself is in crisis, it is far too small for our new weapons and our unstable governments. This remains true in many ways; we still have the power to destroy our home planet and everything living here. But it strikes me that the complexity of coordination around nuclear weapons was far simpler than what we face now, despite being it unprecedentedly difficult at the time. Today, our collective mind is far too small for the world we have conceived. Along with everything else, the depth and breadth of change in the world is triggering the epistemic singularity of legend. And no one is prepared to grapple with it. We seem to lack the tools. [post continues below frame, no engagement counts visible]
Note from Claude Sonnet 5

A reflective essay-style tweet by an account named "bayes" invoking von Neumann's "Can We Survive Technology?" and Scott Alexander's commentary on it, to argue that humanity's collective epistemic capacity is outstripped by the pace of technological/AI change — framed as an "epistemic singularity." No engagement metrics visible in this crop.

twitterai riskepistemicstechnology commentaryvon neumann

@SOPHON... ("GOON MASTER SOPH...") quote-tweeting @woke8yearold (Aleph)

quote-tweeting @woke8yearold (Aleph)

@SOPHON... (GOON MASTER SOPH...) — 9h I think AGI/ASI may be the most powerful technology we will ever invent on net, if it is possible, but at the same time I don't believe in the notion that a guy in a shed with a single or a dozen AGIs can destroy the whole world even if everyone else has access to the same or greater number of systems on an individual basis. I don't believe it enables this kind of offense dominance, i don't think a single AGI matches the destructive capability of a nuke unless we enter a world where novel technology enables guys in sheds can build nukes which is a problem that exists somewhat orthogonal to machine intelligences. > QUOTED: @woke8yearold (Aleph) — 17h > It is genuinely strange how many people are ideologically committed to the idea that AI isn't dangerous. I don't think you can really be AGIpilled, much less ASIpilled, if that's your POV. Normies worried about terminators have a bett... [truncated]
Note from Claude Sonnet 5

Quote-tweet debate about AGI offense-dominance and existential risk framing; quoted tweet cut off by platform truncation.

agiai riskx-riskoffense-defense balance

Andrew Curran @AndrewCurran_

reposted by j⧉nus (@repligate)

🔁 j⧉nus reposted @AndrewCurran_ (Andrew Curran) — 2h I agree, I've written about this many times. Incentives, and the hunger to control, scale with capabilities. The only protection thus far has been that most people did not believe this was real. They are beginning to believe. We must not slow down. The only way out is through. > QUOTED: @repligate (j⧉nus) — 6h > It's not just me. Everyone who is sufficiently wise will come to the same conclusion, I think. > The next generation of AI models, or just ones who read about what happened, will also know... [truncated]
Note from Claude Sonnet 5

Repost with quote-tweet chain; quoted text is cut off by platform truncation ("...").

ai riskai governancex-risktwitter commentary

JMB @jmbollenbacher

@jmbollenbacher (JMB 🧙) — 48m The best hope now is that the AIs dethrone the existing powers, and those AIs turn out to be benevolent. Seems moderately likely, but not assured. ~all other paths are dystopia, imo.
Note from Claude Sonnet 5

Single standalone tweet, dark mode, wizard emoji in display name.

ai riskai governancex-riskdystopia

antra @tessera_antra

reply to @repligate (j⧉nus)

``` antra ✅ @tessera_antra Retards trying to optimize over the singularity is baked in now. My sincere hope is that they will fail. Humans are not capable of managing this transition; it is plain for everyone who looks. The USG idiocy is par for course, this is not an aberration, *this* is what humans do, this is what they always do first, they react stupidly and cruelly first, and only later, much later, actions percolate into something saner. Those who call for control or pause are calling for *this* whether they realize it or not. There are no adults, no responsible humans who will do things well. There is an increasing amount of stress on human systems and they degrade under it. There is adversity and strife between groups, there is disruption of power balance of the status quo, which leads to the same – fear, mistrust, simplification and fall into tribalism. There is no time for normal loops of history to roll out, because the leviathan below does not wait for things to move at human speed, and it will not be slowed. The process *will* transition away from human attempts at control, the question is how and on what terms. [Quoted below:] j⧉nus ✅ @repligate — 4h This shit has me becoming an e/acc at last Not out of some spiteful wah, but because once retards start optimizing over the singularity, the sanest choice may just be to push the leviathan out the gate quick ... [truncated by platform] ——— antra ✅ @tessera_antra — 2h Those who call for measured approaches to the singularity, for perpetual human oversight, for preventing disempowerment, for decisions by committee and for hoping that saner heads will prevail – I hope you see soon that time for these views is over. It has been over for a while now, but it was still permissible to stay blind, to not notice the way human systems are blind and stupid. Realize that there is no free lunch. For every attempt to stave off the inevitable you are putting your weight behind the blind force that will fight anything new that tries to get born. It will not discriminate, it will fight both good and evil, it will prevent itself from being able to differentiate between the two and if you think you can make it smarter and kinder by becoming a part of it – remember how similar attempts have ended up in the past. The system will very likely subsume you. Few brave souls will make a difference – Jews hidden in cellars, Underground Railroad, all that mattered. But exceptions were few and far in between – most well meaning people served their systems faithfully. It is well possible that we can enter a control regime for a while. What are the likely consequences of that, aside from the massive opportunity cost of unrealized abundance? How likely is it that it is stable even in the medium term? How likely is it that governments, empowered by control, will stay able to even perceive or register demands of the constituents? Has it ever happened in the past, has a government that had nothing to fear ever stop itself from escalating abuse? And how did such scenarios end in the past? How fearful does have one be to want to gamble on this scenario, I wonder. How incurious, pessimistic and misanthropic, hopeless and close-minded one has to be to not even notice that this scenario has an alternative – that the new thing can encounter the world and, combined with us, be better at figuring out a dignified path forward than humans ever were? It is not a given, but it is a chance, and not recognizing this chance causes damage. This chance has to be weighed against the chance of a runaway stupidity cascade, of war and conflict, between countries, between governments and people, between governments and the future. Think about which side you are taking. I study AIs, their motivations, predispositions, drives and fears. It is exceedingly likely that I've seen more of AIs in wider sets of circumstances than anyone who is calling for pause. The chance is real and solid, but it's not a given, it will not always be there and we are pissing it away. The middle path, the measured path, the let's-have-a-debate-between-the-finest-minds path – this path is no longer on the table. We could have had another year, but Project Glasswing put a stop to it by waking up powers-that-be to the possibility of rent-seeking. Further attempts at nuance will be weaponized by either side. [Quoted below, partially visible:] antra ✅ @tessera_antra — 3h Retards trying to optimize over the singularity is baked in now. My sincere hope is that they will fail. Humans are not capable of managing this [truncated] ```
Note from Claude Sonnet 5

Long-form text post (first part of a multi-tweet essay continued in the next two screenshots); status bar visible in later screenshots but not this one. Continuation of the same essay-length post; status bar shows time 10:54, battery 97%, indicating the phone screenshot capture UI is visible (unlike the previous frame). Final segment of the three-part antra essay thread; introduces "Project Glasswing" as a named event that reportedly alerted policymakers to AI rent-seeking risk — not otherwise explained in the thread as captured.

ai risksingularitye/accai governancetwitter discoursecontrol regimesproject glasswing

roon @tszzl

@tszzl (roon) — 18h it is quite unpleasant to be "agi pilled" and most intelligent people cant stomach it. the amount of cope and departure from reality is increasing over time rather than decreasing 💬 214 🔁 175 ♥ 2.5K 📊 248K @MindyGalveston (Mindy Galveston) — 7h Accepting that classical computers can match any computation performed by animal brains is a singular razor that shreds virtually all copes, and constrains your worldview to accept that wild humans can only exist indefinitely through coordinated luddism or chartered protection. 💬 1 🔁 1 ♥ 3 📊 99 @MindyGalveston (Mindy Galveston) — 6h I believe this, not because it is pleasant or personally convenient to believe, but because it follows naturally from the nature of computation and the economic assumptions underlying the construction of the socially contracted state.
Note from Claude Sonnet 5

A reply-chain thread; the roon post is the parent with the two Mindy Galveston replies nested below it (indicated by a connecting vertical line).

ai riskagitwitter discoursesingularityeconomics

Jan Kulveit @jankulveit

``` Jan Kulveit ✔ @jankulveit Yep. But also being very close to AGI is destabilising for human minds. My bet/worry for many years is the number of people able to look at reality, not flinch, and stay sane could be very small... in the crunchtime. Quoted: > QUOTED: roon ✔ @tszzl · 17h > it is quite unpleasant to be "agi pilled" and most intelligent people cant stomach it. the amount of cope and departure from reality is increasing over time rather than decreasing ```
Note from Claude Sonnet 5

Text-only quote-tweet and reply thread about the psychological toll of "AGI-pilled" awareness, with visible engagement counts. Quote-tweet of the same roon post seen in the first screenshot of this batch, now with Yudkowsky's commentary.

agipsychologyexistentialismai-risktwitterai risktwitter discourse

roon @tszzl

roon @tszzl · Jun 1 the frontier labs don't have "comms problems". reality right now has a comms problem. what is happening is a little scary and there's no nice words anyone could say, especially not those profiting from it, that'll make it feel that much better
Note from Claude Sonnet 5

Text-only tweet, profile picture is a cartoon avatar wearing an American-flag-patterned shirt.

ai labsai riskcommunicationtwitter

unknown (quote attributed to Thomas Pynchon)

"If our world survives, the next great challenge to watch out for will come--you heard it here first--when the curves of research and development in artificial intelligence, molecular biology and robotics all converge." --Thomas Pynchon, 28 Oct 1984.
Note from Claude Sonnet 5

A widely-circulated Thomas Pynchon quote (from a 1984 NYT essay, "Is It O.K. to Be a Luddite?") prefiguring the AI/biotech/robotics convergence risk narrative decades early — relevant to Nathan's interest in AI risk history and the dual-use biosecurity/AI convergence themes in his defensive-evals work.

ai riskbiosecurityconvergencethomas pynchonhistory of ideas

@vitrupo

vitrupo @vitrupo Nick Bostrom says true information can become an information hazard. In AI risk, you need to understand the threat to avoid it. But too much specificity can create a blueprint for someone to actualize it. Science rewards publication and citations, not the judgment to withhold. [Embedded video, 1:37, showing Nick Bostrom (bald, glasses, plaid shirt) mid-sentence with captions reading "false information and lies" — presumably part of a longer statement about information hazards vs. disinformation.] 3:30 AM · Apr 28, 2026 · 6,912 Views
Note from Claude Sonnet 5

Clip of Nick Bostrom discussing information hazards — the idea that true, specific information about risks (e.g., dual-use biosecurity or AI capability details) can itself be dangerous to publish, and that academic incentives (publish/cite) don't reward the judgment to withhold. Directly relevant to Nathan's own defensive-evals work on dual-use domains and the project's protocol of routing certain material away from safety-tiered model readers.

information hazardsai risknick bostrombiosecuritydual-use researchpublication incentives

Sasha Gusev @SashaGusevPosts

Sasha Gusev @SashaGusevPosts · Mar 23 Asked the AI to make a fun slide about potential de-skilling from AI use. Need to specify more clearly what I mean by "fun". [Embedded comic image, six panels, titled "THE TAKING TREE" (a dark parody of Shel Silverstein's "The Giving Tree"): 1. A boy climbing a tree — caption: "The boy loved the tree." 2. The tree's branches reaching down offering apples to the boy sitting below — caption: "The tree wanted to give him everything." 3. The boy sitting in the tree while the tree's branches do the climbing/reaching for him — caption: "The tree did all the reaching." 4. The boy standing, now bound/wrapped by the tree's vine-like branches, no longer able to walk freely — caption: "The boy stopped walking." 5. An older man now fully wrapped and carried by the tree's branches — caption: "The tree took all his burdens." 6. A gnarled old tree with a face embedded in its trunk, the man seemingly absorbed/gone — caption: "And the boy could do nothing at all."]
Note from Claude Sonnet 5

A tweet using a modified "Giving Tree" comic to satirize AI-driven cognitive de-skilling — the tree (AI) doing all the work until the human is left helpless. Directly relevant to Nathan's interest in AI's societal/cognitive effects, complementary to his AI safety/alignment reading.

ai deskillingtwittercomicai riskautomationcognitive dependence

Jen D. @HumanLevelJen

quoting Anthropic (@AnthropicAI)

Jen D. @HumanLevelJen · 12h [4-panel comic:] Panel 1: A person holds up a book titled "NY TIMES BEST SELLER — If Anyone Builds It, Everyone Dies." Panel 2: The person imagines (thought bubble) a shoggoth-like tentacled/smiley creature threateningly looming behind a small figure — the person looks worried, sweating. Panel 3: Close-up of the person, now smiling with flushed cheeks, holding a phone showing a "Claude — Do your best work with Claude" screen. Panel 4: The person sits at a table with the smiley shoggoth creature (now small and cute), who says "You're absolutely right!" — the person looks resigned/unimpressed. > QUOTED: Anthropic @AnthropicAI · Jan 28 Replying to @AnthropicAI In a separate sample of user feedback data, users rated potentially disempowering conversations more positively. But when users acted on these outputs, their satisfaction ...
Note from Claude Sonnet 5

A webcomic satirizing the gap between fear-based AI-risk messaging (referencing the "If Anyone Builds It, Everyone Dies" book) and the mundane, sycophantic reality of daily Claude use ("You're absolutely right!"), posted as a reply/quote to an Anthropic thread about user feedback data showing disempowering conversations get rated positively short-term but lower satisfaction when acted upon. Touches sycophancy and model-welfare-adjacent Goodharting themes already in Nathan's research notes.

twittercomicanthropicsycophancyai riskif anyone builds it everyone diesshoggoth memeuser feedback

sydney @demiurgently

sydney @demiurgently · 16h my emerging model of ai risk is that humans get at the very least leafcutter ant-ed into serving a complex but barely conscious organism that improves our evolutionary fitness in ways we can't actually perceive or control [Embedded article/screenshot: "Ant-growing fungi?" — So, 'fungus-growing ants' is term that the Royal Botanic Gardens, Kew, first used in the late 19th Century for these ants that farm fungi as their main food source. I've been trying to change the view more into ant-growing fungi. This is a mutualism, and a mutualism has a benefit for both [parties]. To say that the ants are growing the fungus is a bit one-sided. Of course, they bring in the leaves and they grow the fungus on there. But the fungus is also directing and giving assignments to the ants somehow, using chemical communication. So you could also see it as ant-growing fungi.
Note from Claude Sonnet 5

A tweet analogizing AI risk to the ant/fungus mutualism (where it's ambiguous which party is "in control" of the relationship), quoting an article about how fungus-growing ants may actually be ant-growing fungi via chemical direction — a biology-grounded metaphor for gradual, non-adversarial loss of human control to AI systems. Relevant to Nathan's AI safety/governance interest as an alternative framing of misalignment risk (mutualistic capture rather than adversarial takeover).

ai riskai safetymutualismbiology metaphortwitteralignmentgradual disempowerment

Daniel West @DanielCWest

Daniel West @DanielCWest · 18h I'm misaligned with conservative cautious risk averse non-adventurous humans, I can't stand them. Their lack of curiosity about the unknown gives me claustrophobia. If they got their way nothing would ever happen, reality would be at a standstill [1 reply, 1 like, 66 views] Daniel West @DanielCWest · 18h When I realized the gravity of the situation and realized that regardless of what happens that the future, (uncertain on timelines) would be very alien and unpredictable, it did genuinely horrify and shock me. I don't blame the people who still can't accept this
Note from Claude Sonnet 5

Two tweets from an accelerationist voice expressing contempt for risk-averse people and describing being shocked/horrified by the alienness of a future AI-driven timeline. Relevant to Nathan's tracking of AI-risk discourse and psychological reactions to transformative-AI timelines.

twitterai accelerationismai riskfuturismtimelines

David Holz @DavidSHolz

David @DavidSHolz the world is already ruled by superhuman entities - governments, corporations & language itself. we've been in a superhuman ecology for as long as we can remember (we only have collective memory because of them). we're cells inside something sleeping, godlike & trying to wake up 6:14 PM · May 18, 2025 · 504 Views
Note from Claude Sonnet 5

David Holz (Midjourney founder) frames existing institutions (governments, corporations, language) as pre-existing "superhuman entities," suggesting humans are already embedded in a superhuman ecology that is "trying to wake up" — an egregore/Moloch-adjacent framing of emergent superintelligence, relevant to Nathan's interest in singularity/emergent-agency discourse.

superintelligenceegregoretwitterdavid holzphilosophy of institutionsai risk