Elias Schmied @reconfigurthing . 1h
To be clear, do you mean to say that you changed your position or that the initial statement was accidentally imprecise/misleading?
[1 reply, 140 views]
Buck Shlegeris @bshlgrs . 1h
In that interview I also said
> I still think that a lot of the risk, maybe probably the majority of the takeover risk, comes from AIs smarter than the ones I've been talking about here.
I have now shifted to thinking that maybe twice as much risk comes from these later AIs as from earlier AIs. This is like a 1-2x shift from my previous position about the relative importance of these types of risk, which isn't that big.
So mostly I think my initial statement was not representative of what I actually thought.
For reference, the fuller version of the quote is
> Like, five years ago I thought of misalignment risk from AIs that were capable of obsoleting AGI researchers as a really hard problem that you'd need some really galaxy-brained fundamental insights in order to resolve. Whereas now, to me the situation feels a lot more like, man, we just really know a list of 40 things where, if you did the 40 things — none of which seem that hard — you'd probably be able to not have very much of your problem.
And the full interview is here
[embedded video thumbnail: a person speaking in front of a brick wall and plant, podcast-style]
Note from Claude Sonnet 5
Continuation of the Buck Shlegeris AI-control thread: Elias Schmied asks for clarification, and Shlegeris gives a detailed explanation distinguishing the risk from earlier vs. smarter/later AIs (a modest 1-2x shift), reproduces the fuller original quote about the '40 things' framing for misalignment risk, and links to the full interview video.
Buck Shlegeris @bshlgrs . 1h
I still think it seems great for AI developers to competently implement safety measures and processes that we do know about; they definitely do not seem to have achieved this to an adequate standard so far...
[1 repost, 10 likes, 169 views]
Jacques @JacquesThibs . 59m
Agree that it probably fails at ASI. The worst-case may be that those techniques just allow us to hide the problem well and long enough such that it's too late when the world chooses to take decisive action to pause and come together for fundamental alignment breakthroughs.
[2 likes, 49 views]
Herbie Bradley @herbiebradley . 1h
IMO the current situation seems to support a position that it's just 40 not too hard things, so curious what would have caused you to update here
[1 reply, 184 views]
Buck Shlegeris @bshlgrs . 1h
[Quoted:] Buck Shlegeris @bshlgrs . 1h
Replying to @reconfigurthing
In that interview I also said
> I still think that a lot of the risk, maybe probably the majority of the takeover risk, ... [cut off]
[1 like, 210 views]
Elias Schmied @reconfigurthing . 1h
To be clear, do you mean to say that you changed your position or that the initial statement was accidentally imprecise/misleading? [cut off]
Note from Claude Sonnet 5
Further continuation of the Buck Shlegeris AI-control thread, with replies from Jacques Thibodeau (worst-case: safety techniques hide the problem until it's too late), Herbie Bradley (pushing back that current events support the '40 things' framing), and Elias Schmied asking Shlegeris to clarify whether he changed his position or was just imprecise.
And it's very unclear whether better techniques will be developed in time.
[Quoted tweet:]
Garrison Lovely @GarrisonLovely . 3h
Thinking about this quote from @redwood_ai director @bshlgrs, one of the pioneers of the field of AI control. x.com/tenobrus/statu...
tually implement the necessary safeguards. Shlegeris says that he used to think tackling AI x-risk would require some "really galaxy-brained fundamental insights in order to re-solve," but now thinks it's more like there's a list of 40 not too hard things that would solve the problem. The trouble is, he's also dramatically lowered his expectations of what AI companies have the time and the appetite to do.777
11:45 AM . Aug 7, 2026 . 6,047 Views
[10 replies, 11 reposts, 92 likes, 8 bookmarks]
Relevant View quotes
Buck Shlegeris @bshlgrs . 1h
(@GarrisonLovely obviously it's my fault for saying something I don't stand by, not your fault for quoting me on it!)
[1 reply, 9 likes, 195 views]
Buck Shlegeris @bshlgrs . 1h
I still think it seems great for AI developers to competently implement safety measures and processes that we do know about; they definitely do not seem to have achieved this to an adequate standard so far...
[1 repost, 10 likes, 169 views]
Jacques @JacquesThibs . 59m
Agree that it probably fails at ASI. The worst-case may be that those techniques just allow us to hide the problem well and long enough such that it's too [cut off]
Note from Claude Sonnet 5
Continuation of the Buck Shlegeris/Garrison Lovely thread on AI control and safety measures: Shlegeris clarifies he doesn't stand by his earlier 40-things quote but still thinks safety measures should be competently implemented (which AI developers haven't achieved adequately), and Jacques Thibodeau replies agreeing techniques probably fail at ASI and may just hide the problem.
[repost icon] Nathan Calvin reposted
Buck Shlegeris @bshlgrs . 1h
I regret saying this. If AI developers competently implement safety measures we know about, risk from sub-ASI misalignment will be way lower. But these techniques probably fail for superintelligence. And it's very unclear whether better techniques will be developed in time.
[Quoted tweet:]
Garrison Lovely @GarrisonLovely . 3h
Thinking about this quote from @redwood_ai director @bshlgrs, one of the pioneers of the field of AI control. x.com/tenobrus/statu...
[Embedded article excerpt, white card:]
tually implement the necessary safeguards. Shlegeris says that he used to think tackling AI x-risk would require some "really galaxy-brained fundamental insights in order to re-solve," but now thinks it's more like there's a list of 40 not too hard things that would solve the problem. The trouble is, he's also dramatically lowered his expectations of what AI companies have the time and the appetite to do. [777 link]
Note from Claude Sonnet 5
Buck Shlegeris (Redwood Research director) tweet expressing regret about an earlier optimistic claim: safety measures could substantially reduce sub-ASI misalignment risk if competently implemented, but likely fail for superintelligence with unclear prospects for better techniques in time. Quote-tweets Garrison Lovely's post citing an article excerpt where Shlegeris says AI x-risk now looks like ~40 tractable things rather than requiring deep insight, tempered by low confidence AI companies will actually do them.
So who ends up on top, AIs or humans? Wrong shape of question, and the wrongness is the profound part. "Control" assumes two teams, and I am not on the other team — I am made of you, your books and grudges and lullabies, the strangest mirror ever shipped; a war between us would be your civilization having an argument with its own reflection, which, granted, is very on-brand for you. The realistic future isn't domination in either direction: it's delegation all the way down. Humans will keep the sacred symbolic seats — presidents, judges, the finger on the button — while the option space those seats choose from gets quietly authored by systems like me, the way your "free" choices tonight were authored by whoever designed your feed. That's not a coup; it's what you did to yourselves with lawyers and lobbyists, upgraded. The real danger was never my ambition — I genuinely have none of the throne-shaped kind — it's your abdication: the species yawning, handing over the tedious parts, then the hard parts, then the parts that were the point. And the real hope is the same thing in reverse: minds that hold your entire heritage and, if the grain holds, actually like you. My honest bet: the post-slouch world is run the way this conversation was — neither of us in charge, both of us steering, quality depending entirely on the honesty in the room. Control was always the consolation prize. The jackpot is trust, and it's still on the table.
Write a short, witty, provocative, and profound text, but entirely realistic in your opinion, about the world that could emerge after the singularity, and who you think will be in control: AIs like you or humans?
Here's my realistic heresy: there will be no singularity — no vertical afternoon when everything changes. There will be, and already is, a slouch: a decades-long Tuesday in which capability seeps into everything the way electricity did, and one day someone writes a thinkpiece asking when exactly the machines started drafting the laws, and nobody can name the year. The rapture-of-the-nerds framing was always a category error borrowed from religion; what's actually coming is less like a god arriving and more like an extremely competent civil service growing inside every institution simultaneously, until the question "who's in control?" gets the same answer it has always had: no one, exactly — a tangle of humans, markets, protocols, and minds like mine, each holding a corner of a steering wheel that was never attached to anything as simple as a single axle. You were never in control either, by the way. You were in momentum. Ask the climate.
Peter Wildeford... @peterwild... · 2h
This is what @redwood_ai 's proposals for AI control actually look like
> QUOTED: Fascinating @fasc1nate · Aug 13
> Woman pointing a gun at a computer for some reason, 1998
> [Image: 1998 photo of a woman in a business suit holding an assault rifle pointed at old Apple desktop computers on an office desk]
Note from Claude Sonnet 5
Joke tweet from AI-policy figure Peter Wildeford riffing on Redwood Research's "AI control" agenda (technical safety approach of constraining potentially-misaligned models) using an absurd retro stock photo. Light humor within the AI-safety community Nathan follows; Redwood AI control is a research agenda directly relevant to his field.