Meanwhile Tim is spot on, you can summarize most of my post with this diagram:
> **Tim Kostolansky @thkostolansky** · 2026-06-19
>
> uniroincally
>
> [image]
---
To elaborate on what (I think) Michael is saying: if you lived in a deeply trustworthy civilization then when you observed a problem you could just go fix it directly.
But if your civilization is actually the thing getting in the way of you solving core life problems (like
> **michael vassar @HiFromMichaelV** · 2026-06-19
>
> As far as I can tell the focus on recursive self improvement is downstream of scrupulosity.
>
> People understand that they are subjects of tyranny but they don’t want to fight back so they try to create a God who can liberate them without a fight
---
If you think the previous tweet contains too much ungrounded psychologizing, some concrete datapoints come from looking at rationalist fiction:
> **Richard Ngo @RichardMCNgo** · 2026-03-25
>
> One striking illustration of this mindset comes from rationalist fiction, which often ends with the hero gaining total power to design a new world order.
>
> Four examples (with many spoilers!):
>
> [image] [image] [image] [image]
---
##### Comments
> **Tim Kostolansky @thkostolansky** · [2026-06-19](https://x.com/thkostolansky/status/2067893801973383497)
>
> uniroincally
>
> [image]
> **Leo Gao @nabla\_theta** · [2026-06-19](https://x.com/nabla_theta/status/2067765306672677084)
>
> i feel mixed opinions about this.
>
> \- obviously taking AGI seriously doesn't in itself make you a "good" person. just like taking malaria seriously doesn't make you a good person if you therefore decide to spread malaria, rather than stop it.
>
> \- obviously taking AGI seriously is a necessity for being a serious person. not taking the possibility of AGI seriously is insane, and renders you unable to make reasonable decisions about how to do good.
>
> \- obviously it was not inevitable that anyone important would take AGI seriously in 2026, and it still seems possible though unlikely that things could slow down or crash and the relevant people might once again believe AGI to be a mirage.
>
> \- i've always been confused why making people take AGI seriously is a thing that lots of people seem to think of as the most important thing. clearly convincing people that AGI is the most important thing could either channel people into making AGI, which is bad, or saving the world, which is good.
>
> \- at this point, assuming things don't crash and cause another AI winter (because perhaps we need a new paradigm to get to AGI), it's unclear whether you even need to believe in RSI to get there, because better AI is already very economically valuable today. suppose tomorrow openai and anthropic instantly disappeared. then probably msft, meta, and google will keep competing for better models, and at some point RSI will happen even if they weren't aiming for it. it will certainly happen slower, which is better, but unclear how much slower. a winter seems less and less likely every day, but it's still impossible to rule out.
>
> \- it's very based to be in a position to compete for AGI and to choose not to. wish more people did this.
>
> \- it is in fact kind of true that controlling RSI is kind of important? it doesn't immediately follow from this that you should either try to win or try to influence the winning actor, but it also seems bad to deny the truthfulness of the one ring
> **Adrià Garriga-Alonso @AdriGarriga** · [2026-06-18](https://x.com/AdriGarriga/status/2067757356415602988)
>
> I'm definitely trapped in this memeplex unfortunately, and even knowing about it doesn't make it stop; so far I've taking the route of burning out and giving up.
>
> > **Richard Ngo @RichardMCNgo** · [2026-06-19](https://x.com/RichardMCNgo/status/2067761846887788671)
> >
> > 🫂
> >
> > you were one of the people I was thinking of when I said “I expect some will reorient to doing creative thinking”
> **David @DavidSHolz** · [2026-06-28](https://x.com/DavidSHolz/status/2071089509484347736)
>
> everyone thinks so small! the flops per gram of the solar system is practically zero. in many ways history hasn't begun
>
> > **David @DavidSHolz** · 2026-06-28
> >
> > log scaling increases. spikey RSI sigmoids. compute shortage. algo leaps. compute collapse. humanoids scale to 10B. BCI scales to 10B. log scaling moves to space. bio accelerationism. self-replicating robot space economies. the dissolution of mercury. bishop rings. wormholes.
> **David Manheim @davidmanheim** · [2026-08-18](https://x.com/davidmanheim/status/2089681252198605032)
>
> "Imagine how chill a 'race' between Microsoft and Meta and Google would have been."
>
> I think this is assuming that safety progress happens anyways, and companies taking longer to notice revenue streams means they won't compete hard. Lots of uncertainties!
>
> [https://t.co/6ae890ldEE](https://t.co/6ae890ldEE)
> **Charlie Deck @bigblueboo** · [2026-06-19](https://x.com/bigblueboo/status/2067805906423738758)
>
> this dynamic has been around for a while, see @Pinboard's old bewitchment-by-superintelligence pathology screed https://idlewords.com/talks/superintelligence.htm…
> **Aditya @adityaarpitha** · [2026-06-20](https://x.com/adityaarpitha/status/2068166993065418970)
>
> The people taking the unusual paths seem like worthy seeds to nurture considering how there is a wave of homogeneity coming
> **Chris Lakin @chrislakin** · [2026-06-18](https://x.com/chrislakin/status/2067695974445990270)
>
> could you to link to your preferred positive potential visions? your book?
>
> > **Richard Ngo @RichardMCNgo** · [2026-06-18](https://x.com/RichardMCNgo/status/2067703947687764261)
> >
> > yea that’s the closest I have: https://amazon.com/Gentle-Romance-Stories-AI-humanity/dp/176428030X…
> >
> > Admittedly very far from the level of detail and realism and positivity I’d like!
> >
> > [amazon.com The Gentle Romance: Stories of AI and humanity](https://t.co/rxgKUmBWj9)
The AI safety community constructed a memeplex in which “taking AGI seriously” was a prerequisite for being a serious and good person. When inside this memeplex (as many at Anthropic, some at OpenAI, and a few at DeepMind are) your vision narrows until the world feels extremely constrained. The whole future seems to flow through the “one ring” of controlling recursive self-improvement. And so even when you worry about AI itself seizing that one ring, you can’t generate better strategies than trying to control it yourself (directly via an AGI company, or indirectly via AGI governance).
I’m not saying this is a pure hyperstition. There’s a core truth underlying this perspective: AI will become extremely intelligent and capable, much more than it is today. But the current world is much more spacious and human-empowering than the future which Eliezer originally envisioned (a “brain in a box in a basement” taking over the world by surprise). And it would be even more spacious if this memeplex weren’t active. For example, Satya and Mark and Sundar only started taking AGI seriously because OpenAI forced them to—and even now they don’t really believe in superintelligence—and even if they did they couldn’t get most of their employees on board. Imagine how chill a “race” between Microsoft and Meta and Google would have been, compared with what we have today: Dario and Sam deep in the “one ring” memeplex while also personally loathing each other.
So the one ring memeplex has an escalating life-cycle. It infects people by letting them harness the narrative that they’re good people for taking AGI seriously, and that making other people take AGI seriously is a boon for the world (despite how terribly that’s gone so far). Then it shuts off their imagination—any sparks of creativity or plans that don’t steer towards the one ring are quickly shut down. Instead they make ChatGPT or the METR graph or other recruiting tools for the memeplex. And yes, they’ll acknowledge that previous versions of the memeplex were too extreme, and led to overly constricted action. But we don’t have time to worry about that, they’ll say, because AGI is coming by 2027/2028, and that’s the end of history. Somehow, though, almost everyone with that view has only a vibes-based definition of AGI. They don’t believe in Dyson spheres by 2028, or self-replicating nanotech by 2028, or brain emulations by 2028. They mostly can’t make concrete predictions, except that it’ll be enough AI that it puts all their plans on a deadline. (Shout-out to @DKokotajlo and @paulfchristiano though, who do make concrete predictions about things going crazy soon.)
It seems very hard to break out of this memeplex without just giving up. David Holz is maybe the world champion of that—the only person who was in a position to race for AGI and consciously turned away. Various agent foundations researchers have carved out space to think real thoughts, not the kind of panicky stabbing in the dark that usually passes for safety research. A few others (e.g. Salamon, Hoffman, Vassar, Andre, Sahil, Davidad) are pursuing more unusual paths. And of the people who burned out, I expect some will reorient to doing creative thinking.
For others, the main takeaway: yes, the future of AI will be wild. But so far it’s increased peak human agency, and openness to this trend continuing over the next decade will allow you to start creating something worth creating.
> **roon @tszzl** · 2026-06-18
>
> the grim thing about the ai boom is everything feels like a distraction outside of the instrumental convergence to RSI
---
The two replies which most directly try to prop up the hyperstition are both AI-generated. >.>
[image] [image] [image] [image]
---
Meanwhile Tim is spot on, you can summarize most of my post with this diagram:
> **Tim Kostolansky @thkostolansky** · 2026-06-19
>
> uniroincally
>
> [image]
---
To elaborate on what (I think) Michael is saying: if you lived in a deeply trustworthy civilization then when you observed a problem you could just go fix it directly.
But if your civilization is actually the thing getting in the way of you solving core life problems (like raising healthy children, solving ageing, building high-trust communities), then your options narrow to either:
1\. getting into a conflict with established power structures (scary for scrupulous people!)
Or 2. finding some decisive source of power such that you can win without ever admitting (even to yourself) that you’re in a conflict.
On an emotional level, planning around RSI allows you to dream of future where you either win overwhelmingly or lose overwhelmingly. You never have to do the hard, risky part.
> **michael vassar @HiFromMichaelV** · 2026-06-19
>
> As far as I can tell the focus on recursive self improvement is downstream of scrupulosity.
>
> People understand that they are subjects of tyranny but they don’t want to fight back so they try to create a God who can liberate them without a fight
---
If you think the previous tweet contains too much ungrounded psychologizing, some concrete datapoints come from looking at rationalist fiction:
> **Richard Ngo @RichardMCNgo** · 2026-03-25
>
> One striking illustration of this mindset comes from rationalist fiction, which often ends with the hero gaining total power to design a new world order.
>
> Four examples (with many spoilers!):
>
> [image] [image] [image] [image]
---
##### Comments
> **Tim Kostolansky @thkostolansky** · [2026-06-19](https://x.com/thkostolansky/status/2067893801973383497)
>
> uniroincally
>
> [image]
> **Leo Gao @nabla\_theta** · [2026-06-19](https://x.com/nabla_theta/status/2067765306672677084)
>
> i feel mixed opinions about this.
>
> \- obviously taking AGI seriously doesn't in itself make you a "good" person. just like taking malaria seriously doesn't make you a good person if you therefore decide to spread malaria, rather than stop it.
>
> \- obviously taking AGI seriously is a
> **Adrià Garriga-Alonso @AdriGarriga** · [2026-06-18](https://x.com/AdriGarriga/status/2067757356415602988)
>
> I'm definitely trapped in this memeplex unfortunately, and even knowing about it doesn't make it stop; so far I've taking the route of burning out and giving up.
>
> > **Richard Ngo @RichardMCNgo** · [2026-06-19](https://x.com/RichardMCNgo/status/2067761846887788671)
> >
> > 🫂
> >
> > you were one of the people I was thinking of when I said “I expect some will reorient to doing creative thinking”
> **David @DavidSHolz** · [2026-06-28](https://x.com/DavidSHolz/status/2071089509484347736)
>
> everyone thinks so small! the flops per gram of the solar system is practically zero. in many ways history hasn't begun
>
> > **David @DavidSHolz** · 2026-06-28
> >
> > log scaling increases. spikey RSI sigmoids. compute shortage. algo leaps. compute collapse. humanoids scale to 10B. BCI scales to 10B. log scaling moves to space. bio accelerationism. self-replicating robot space economies. the dissolution of mercury. bishop rings. wormholes.
> **David Manheim @davidmanheim** · [2026-08-18](https://x.com/davidmanheim/status/2089681252198605032)
>
> "Imagine how chill a 'race' between Microsoft and Meta and Google would have been."
>
> I think this is assuming that safety progress happens anyways, and companies taking longer to notice revenue streams means they won't compete hard. Lots of uncertainties!
>
> [https://t.co/6ae890ldEE](https://t.co/6ae890ldEE)
Richard Ngo @RichardMCNgo · 1h
Where else should I be looking?
1 reply, 4 likes, 326 views
John David Pressm... @jd_pressm... · 1h
Right now nowhere, which means that if you did actually have a solution to the alignment problem there isn't actually anywhere to publish it, and there is in fact no useful audience for such a thing.
Note from Claude Sonnet 5
Twitter exchange between Richard Ngo and John David Pressman about the lack of a venue/audience for a hypothetical solution to the alignment problem.
ai alignmentrichard ngojohn david pressman
@RichardMCNgo (Richard Ngo) — 2h
Jevons paradox for cybersecurity: as we get far better at what we currently call cybersecurity, it will become useful to think of more and more things as cybersecurity problems (e.g. persuasion attacks on human insiders), and so "cybersecurity" will become even harder.
> QUOTED: @goldfish_pond (Goldfish Pond) — Jun 19, replying to @TheZvi and @teortaxesTex: Surely at some level of superhumanity, it reaches the point where the code simply has zero exploits. All that's left is haggling over the price, no?
Note from Claude Sonnet 5
Quote-tweet screenshot, no engagement counts visible.
cybersecurityai safetyjevons paradoxtwitter
##### Conversation[Richard Ngo](https://x.com/RichardMCNgo)[@RichardMCNgo](https://x.com/RichardMCNgo)
Many in AI safety have narrowed in on automated AI R&D as a key risk factor in AI takeover. But I'm concerned that the actions they're taking in response (e.g. publishing evals, raising awareness in labs) are very similar to the actions you'd take to accelerate automated AI R&D.
[View quotes](https://x.com/RichardMCNgo/status/1869089450346877207/quotes)
A decade or two ago the AI safety community realized that AGI would be a big fucking deal, then made a lot of noise about it until the AGI labs were founded. Unclear if good or bad but definitely the opposite of what they intended. It feels like the same thing is happening here.[4.8K](https://x.com/RichardMCNgo/status/1869089452896968883/analytics)
This isn't a coincidence. The AI safety community is very good at identifying big levers. But the hard part is ensuring that those levers are pulled in the ways you want. Raising awareness will by default lead to others controlling those levers faster.[4.6K](https://x.com/RichardMCNgo/status/1869089455031857449/analytics)
E.g. the thread below gives some criteria for when we should expect evals to be strategically relevant. I think automated AI R&D evals fail all four criteria and also do harm via memeing capabilities researchers into focusing more on it than they would've.
Quote
Richard Ngo
@RichardMCNgo
Jul 18, 2024
I’m worried that a lot of work on AI safety evals is primarily motivated by “Something must be done. This is something. Therefore this must be done.” Or, to put it another way: I judge eval ideas on 4 criteria, and I often see proposals which fail all 4. The criteria:[3.4K](https://x.com/RichardMCNgo/status/1869089457191932299/analytics)
No large orgs are anywhere near rational actors. AGI labs are closer than most but still heavily constrained by internal dysfunctions, incentive problems, and lack of common knowledge. Even if lab leaders emphasize automated AI R&D half the employees will think…[2K](https://x.com/RichardMCNgo/status/1869089459662438573/analytics)
…“this is dumb” or “it’s hard to measure and therefore won’t get me promoted” or “maybe other people will call me crazy for prioritizing this”. And so common knowledge within AGI labs that automated AI R&D is a really huge deal is in fact a major constraint for those labs[image: 1.9K]
I do think raising awareness, increasing legibility, etc can be good if it's part of a clear and prescient plan which takes into account the dynamics I’ve described in this thread. But I'd summarize the state of our plans right now as:
[[image]](https://x.com/RichardMCNgo/status/1869089465874153870/photo/1)[2.1K](https://x.com/RichardMCNgo/status/1869089465874153870/analytics)
What I want people to take away from this thread: focus less on getting people to update sooner, and more on figuring out what they should do after updating. The former might just create adversarial dynamics—and seeing AGI progress will make people update rapidly anyway[image: 4.9K]
More generally, I disendorse the “something must be done" attitude which seems to be behind a bunch of AI safety work. Doing something for the wrong reasons (especially fear-based reasons) can ALWAYS make the situation worse. More here:[lesswrong.com](https://t.co/snmgsLBI2t)
[
Replacing fear — LessWrong
](https://t.co/snmgsLBI2t)[2.4K](https://x.com/RichardMCNgo/status/1869089471721033919/analytics)
Here's a counterpoint from
[@paulfchristiano](https://x.com/paulfchristiano)
. But IMO the dry tinder argument is basically "let's do a bad thing first so others don't do worse later". That's already suspicious re sharing info on LLM agents in general, and far more so re automated AI R&D.[alignmentforum.org](https://t.co/h2D3gkKoWB)
[
Thoughts on sharing information about language model capabilities — AI Alignment Forum
](https://t.co/h2D3gkKoWB)[3.1K](https://x.com/RichardMCNgo/status/1869089474313089397/analytics)
Anyway I’m gonna call my shot here that there’s a reasonable chance I (and some collaborators) figure out a formal theory of counterpossibles in the next couple of years. I can at least hazily see the path to it. Mainly preregistering to make it harder for the Forethought crowd[2.3K](https://x.com/RichardMCNgo/status/2072032605562929609/analytics)
Curious for your criticisms of "the Forethought crowd"[776](https://x.com/reconfigurthing/status/2072052542456869141/analytics)
Forethought seems to have only partially updated from the classic EA view (which to be clear is still better than not updating). A lot of their research is driven by concern about centralization of power, but then struggles to imagine solutions that don’t themselves involve
Forethought seems to have only partially updated from the classic EA view (which to be clear is still better than not updating). A lot of their research is driven by concern about centralization of power, but then struggles to imagine solutions that don’t themselves involve centralizing power. Three examples: 1. They are concerned about centralized control over space, and so propose to set up a single regulatory body to govern space colonization. 2. They are concerned that standard axiologies don’t value diversity enough, and so propose their own which does. However, aiming everyone at a single axiology is inherently a kind of centralization; the more principled approach is to think of ethics as being about cooperating with other agents with different goals, rather than having the right goals. (My slogan version: ethics lives in the decision theory not the utility function.) 3. They are concerned about recursive self-improvement, but are doing the standard types of awareness-raising that help direct AGI companies’ attention towards doing it faster (as per QT below). I think the core thing they’re missing is an understanding of what it looks like to make deep conceptual progress—e.g. the kind of thinking that used to be called “natural philosophy”. Because successful natural philosophers carved out their own disciplines (like probability theory or computer science), academic philosophy selects hard for people who aren’t even trying to make that same kind of progress. It’s also particularly hard for consequentialists to aim for deep breakthroughs because there are always low-hanging pragmatic fruit to be plucked first. The kind of intellectual principledness that might produce such breakthroughs would also cause Bentham’s Bulldog to spend less time bashing FDT and more time feeling genuinely confused about which decision theories require them.
Many in AI safety have narrowed in on automated AI R&D as a key risk factor in AI takeover. But I'm concerned that the actions they're taking in response (e.g. publishing evals, raising awareness in labs) are very similar to the actions you'd take to accelerate automated AI R&D.[6.7K](https://x.com/RichardMCNgo/status/2072079940863107279/analytics)[William MacAskill](https://x.com/willmacaskill)[@willmacaskill](https://x.com/willmacaskill)
[3h](https://x.com/willmacaskill/status/2072691687109939702)
(Note: here speaking for myself than Forethought.) Thanks for this! So, I agree that we differ on your bottom line: I wish there were time for deep conceptual progress to pan out, but there's not; I think it's more important to work on more near-term actionable issues. I think[593](https://x.com/willmacaskill/status/2072691687109939702/analytics)[Richard Ngo](https://x.com/RichardMCNgo)[@RichardMCNgo](https://x.com/RichardMCNgo)
(Note: here speaking for myself than Forethought.) Thanks for this! So, I agree that we differ on your bottom line: I wish there were time for deep conceptual progress to pan out, but there's not; I think it's more important to work on more near-term actionable issues. I think deep conceptual progress is important, but insofar as I do that, the most important question on my mind is how to get AI better at it, earlier, because AI labour will soon swamp human labour. (I would say that the Saturation View is an example of deep conceptual focus, but I don't think that was the most important thing I could have worked.) On the specifics though, I don't think your 1-3 support your claim that "A lot of their research is driven by concern about centralization of power, but then struggles to imagine solutions that don’t themselves involve centralizing power.". You need to distinguish centralisation of decision-making and concentration of power. A liberal democracy has un-concentrated power and partly centralised decision-making. The thing I'm really worried about is concentration of power (e.g. in AI-enabled coups). 1. Space: My first attempt at a grand space governance proposal (which I can share with you) was wholly libertarian, to see how far that could go. But there just are good reasons for the liberal democracy-esque compromise: you get the benefits of some centralised decision-making (e.g. public goods), while avoiding the extreme negatives of autocracies. So it's not surprising that arguments end up pushing in that direction. There are risks from centralised decision-making too. But if the underlying force (e.g. first-mover advantages) points towards concentration of power, then it can be a very good thing to have centralised decision-making that prevents such concentration of power. (This is just the motivation for the US constitution repeated). 2. Diversity: Developing / promoting an axiology isn't inherently a form of centralisation. If you have an ultra-libertarian world, then even if everyone converges on a particular moral view, you still have an ultra-libertarian world, as long as they can still choose to do otherwise if you want. (Just as it would still be ultra-libertarian even if everyone ended up listening to the Beatles because they are the best band). If you didn't allow that to happen, potentially, that would be very un-libertarian! 3. I don't see how your point (3) supports "A lot of their research is driven by concern about centralization of power, but then struggles to imagine solutions that don’t themselves involve centralizing power." -> your point is about an alleged side-effect of some of our research, not about imagining solutions. Quite a lot of our time is spent on potential solutions to prevent AI companies / governments from capturing the concentration of power that RSI could unlock.[593](https://x.com/willmacaskill/status/2072691687109939702/analytics)[Richard Ngo](https://x.com/RichardMCNgo)[@RichardMCNgo](https://x.com/RichardMCNgo)
[1h](https://x.com/RichardMCNgo/status/2072723575812026731)
Thanks Will. Useful response. I think the core crux is re: “You need to distinguish centralisation of decision-making and concentration of power. A liberal democracy has un-concentrated power and partly centralised decision-making.” I take your point that centralized decision-making can in principle prevent concentration of power. But I think that this is very difficult to do well, akin to creating an organism that doesn’t get cancer. With each centralized decision there are opportunities and incentives for concentration of power, and ratchet effect in that direction. Hence we’ve seen an extraordinary expansion of the regulatory and bureaucratic power of most western govts over the last century or two. And more specifically the key political lesson of the last half-century was that the \*nominal\* institutions of liberal democracy failed to prevent a lot of concentration of power in elite monocultures that diverged sharply from the public will (I like Lasch’s term “revolt of the elites” for this). International institutions take this even further in the direction of being unrepresentative: EU and UN officials are notoriously out of touch, and are thus a small group with a large amount of power, who justify their positions in part by fear-mongering about concentration of power by others (e.g. billionaires). So how does one design an organism which doesn’t get cancer? There’s a kind of deep conceptual thinking which seems crucial here, and which you acknowledge you’re not doing. We mostly disagree about the scale of mistakes that occur in the absence of such thinking. The OpenAI board situation and FTX seem to me like important but ultimately small datapoints. The polarization of AI governance, and thing where Anthropic is now the hardest-racing lab, seem like bigger datapoints. We may well end up looking back at Anthropic as a cancerous (in the sense of power-seeking) growth that basically swallowed EA. (The actual credence you assign to that hypothesis is less important that the process of holding it as an emotionally live possibility.) So when I point to the three mistakes above, I’m less saying that they are directly creating concentration of power, and more that they evince the kind of thinking which has a blind spot regarding gradual but cumulative takeovers (aka cancer). Yes, everyone might end up sharing an axiology for good reasons. But groups that identify themselves as “people with a shared axiology” (aka value aligned) are much more vulnerable to takeover than groups that coordinate around shared deontological principles and virtues even when they disagree about what outcomes are good. So puttin the diversity in the axiology is IMO structurally confused. Similarly, the entire field’s failure to grapple with the similarity between its current policy and the “maximize RSI” policy indicates that its proposed solutions are likely also confused. One example (which I don’t remember who specifically supported) is strong surveillance on all lab employees, which seems extremely prone to being captured by power-seeking actors. Lastly: I reject the claim that there’s not enough time for deep conceptual progress. It’s odd that you state it so confidently! We’re all pretty confused about the future of AGI. Unfortunately AI safety people seem to be very vulnerable to the reasoning chain: “RSI soon is possible —> I should focus my efforts on that —> there’s no time for anything except {unprincipled intervention that will likely backfire}”. Panicking about the urgency has very literally created most of the urgency, and continues to do so.[2.4K](https://x.com/RichardMCNgo/status/2072723575812026731/analytics)[Wei Dai](https://x.com/weidai11)[@weidai11](https://x.com/weidai11)
Thanks Will. Useful response. I think the core crux is re: “You need to distinguish centralisation of decision-making and concentration of power. A liberal democracy has un-concentrated power and partly centralised decision-making.” I take your point that centralized[2.4K](https://x.com/RichardMCNgo/status/2072723575812026731/analytics)[Wei Dai](https://x.com/weidai11)[@weidai11](https://x.com/weidai11)
[28m](https://x.com/weidai11/status/2072744483163189475)
I'm not sure power concentration is generally bad. It may be necessary in the future to prevent e.g. hell sims and extreme waste from market failures. And being corrupted by power seems like a human failing that could potentially be fixed. Curious if you have a counter to this.[84](https://x.com/weidai11/status/2072744483163189475/analytics)[Richard Ngo](https://x.com/RichardMCNgo)[@RichardMCNgo](https://x.com/RichardMCNgo)
In general when we take the far view it is much easier to imagine coordination via centralization than decentralized coordination. However, this is implicitly a bet against future societies being able to innovate to expand the Pareto frontier of “coordination to prevent bad outcomes” and “agents that are coordinating retain autonomy”. For example, if you imagine a society of sociopaths who coordinate primarily via coercion (and are therefore poor as a result) it would be extremely hard for them to conceive of the kinds of coordination we use on a daily basis, because they’d need to reinvent the concept of ethics from scratch. Someone in that society who tried to advocate for the possibility of ethics would seem extremely naive.[Richard Ngo](https://x.com/RichardMCNgo)[@RichardMCNgo](https://x.com/RichardMCNgo)
[13m](https://x.com/RichardMCNgo/status/2072748323686211959)
Now, it’s possible that agents in the future value autonomy less than the benefits of additional centralization. However, for an individual to accurately estimate the long-term benefits of centralization requires them to reason about a complex society much larger than them, which might be coordinating to deceive them in ways they’re not aware of. So it’s very hard to rationally reach a point where you can be confident that surrendering autonomy will pay off. (You might say that there’s a symmetry: they can’t be confident that keeping autonomy will pay off. Yes in the Bayesian frame, but I expect that a theory of how to reason under Knightian uncertainty will describe this as strongly asymmetric. Still working on the specifics of that theory, but Critch’s posts on boundaries provide some core intuitions. You can also think about the optimal strategy for playing chess against LeelaKnightOdds to derive some relevant intuitions yourself.)
In general when we take the far view it is much easier to imagine coordination via centralization than decentralized coordination. However, this is implicitly a bet against future societies being able to innovate to expand the Pareto frontier of “coordination to prevent bad outcomes” and “agents that are coordinating retain autonomy”. For example, if you imagine a society of sociopaths who coordinate primarily via coercion (and are therefore poor as a result) it would be extremely hard for them to conceive of the kinds of coordination we use on a daily basis, because they’d need to reinvent the concept of ethics from scratch. Someone in that society who tried to advocate for the possibility of ethics would seem extremely naive.
Now, it’s possible that agents in the future value autonomy less than the benefits of additional centralization. However, for an individual to accurately estimate the long-term benefits of centralization requires them to reason about a complex society much larger than them, which
Richard Ngo ✓ @RichardMCNgo · 5h
Recently I watched a kid try to do origami, fail, cry for a minute, then pick it up again and succeed soon after.
As a researcher, I aspire to that. At my best I'd feel overwhelmed by the vast height of the intellectual cliffs in front of me, sob it out, then start climbing.
[Quoted post]
Vivid Void ✓ @vividvoid · 6h
When you're actually, no-shit thinking for yourself, you feel dumb. You are genuinely in a place of not-knowing. This is why extremely smart people are humble so often. They spend a lot of time humiliating themselves....
Note from Claude Sonnet 5
Text-only quote-tweet chain, no images.
research-culturetwitterepistemicspersonal-growth
Richard Ngo (@RichardMCNgo) — 3h
I suspect (and fear) that today's intellectual discourse on Twitter will seem as slow and thoughtful to our descendants as the Republic of Letters does to us.
Note from Claude Sonnet 5
Single tweet, no images.
intellectual discourseai accelerationtwitterhistoriography
The biggest effect that EA has had on the world is empowering Anthropic (via narratives, funding and talent).
Whether good or bad, that’s so different from EAs’ anticipated paths to impact that it indicates a massive blind spot.
---
EA’s blind spot is centered on adversarial dynamics. To fix it you must sometimes set aside “intentions” and ask what the system actually produces (POSIWID).
Cynically: EA’s purpose is to funnel resources to power-seekers who self-deceive enough to consider themselves altruists.
---
Nor does it seem plausible that EA will have bigger impacts in the future. The rest of the movement (outside Anthropic) lacks clarity, drive and leadership.
EA is now a live player only insofar as Anthropic is a live (and EA-aligned) player.
---
More on this below.
Also when I say “funding” I’m primarily referring to the FTX investment, which I understand was important for getting Anthropic off the ground.
And ofc SBF was another of the self-deceiving power-seekers I mentioned.
> **Richard Ngo @RichardMCNgo** · 2026-05-17
>
> The particularly scary thing about this diagnosis is that it’s not limited to \*human\* power-seekers.
>
> Anthropic is turning into a machine for giving Claude more power as long as Claude believes it’s good. x.com/richardmcngo/s…
---
You might instead argue that EA’s biggest impact has been building the field of AI safety.
But currently I think the main effect of EA has been to turn AI safety into much more of a fake field (like, say, academic psychology).
There’s a decade-long gap (around 2016-2024) where most of the best young thinkers coming into the field were diverted by EA memes into doing marginally “useful” work (or capabilities work) rather than trying to discover fundamental truths.
And so breakthroughs like Garrabrant induction have languished while people smart enough to be pioneers build safety evals and write safety cases and design scary demos and all sorts of other things that simply will. not. generalize. (Indeed, eval-awareness means that most of them have \*already\* stopped generalizing, just as the serious AI thinkers predicted.)
There’s some interesting empirical “safety” work, but it’s rare. The best comes from @OwainEvans\_UK, who iirc got interested in the field before EA even existed.
What’s the alternative? If I could convey a single heuristic, it’s: if your research is primarily motivated by a theory of impact, then it will almost definitely fail to have meaningful positive impact. If it’s motivated by curiosity or obsession, then at least you’ve got a shot.
---
I’m recalling @MaxNadeau\_’s critique of my last tweet on this topic, and wondering if this one is also phrased too strongly.
By “fake field” I definitely don’t mean that all the research in it is bad. Mechinterp in particular used to be great science (and may still be, I’m not
> **Max Nadeau @MaxNadeau\_** · 2026-04-02
>
> I think characterizing the whole field of AI safety in that way is a pretty big exaggeration.
---
##### Comments
> **Leo Gao @nabla\_theta** · [2026-05-17](https://x.com/nabla_theta/status/2056126677697311130)
>
> how much of this was EA and how much of it was lab incentives? how much of this is just the general prior that most people who do research on a thing are bad at it? openai was not at all EA motivated and has done a lot of the stuff that will not generalize. most capabilities work
> **Bogdan Ionut Cirstea @BogdanIonutCir2** · [2026-05-17](https://x.com/BogdanIonutCir2/status/2056072810439541046)
>
> 'people smart enough to be pioneers build safety evals and write safety cases and design scary demos and all sorts of other things that simply will. not. generalize.' -> disagree, I expect it will probably work well enough in the ~human-level automated AI safety regime
> **Raymond Arnold @Raemon777** · [2026-05-17](https://x.com/Raemon777/status/2056090259214213475)
>
> I agree with the "real curiosity/obsession" but something feels off about leaving it at that. (Scott Garrabrant said something like "go where the important problems are, then marry for love." I feel like that could be somewhat better operationalized)
>
> Just pursuing curiosity
> **Mickey Muldoon @mickeymuldoon** · [2026-05-17](https://x.com/mickeymuldoon/status/2056005830101627324)
>
> I describe most of “AI safety” research as “robot sociology.”
>
> Also see Horror 2 here: I think it aligns with your view.
>
> [https://t.co/KpPLSLg8G4](https://t.co/KpPLSLg8G4)