← All topics

redwood research

4 captures, most recent first.

Buck Shlegeris @bshlgrs

quoting @GarrisonLovely — saved image

[repost icon] Nathan Calvin reposted

Buck Shlegeris @bshlgrs . 1h
I regret saying this. If AI developers competently implement safety measures we know about, risk from sub-ASI misalignment will be way lower. But these techniques probably fail for superintelligence. And it's very unclear whether better techniques will be developed in time.

[Quoted tweet:]
Garrison Lovely @GarrisonLovely . 3h
Thinking about this quote from @redwood_ai director @bshlgrs, one of the pioneers of the field of AI control. x.com/tenobrus/statu...

[Embedded article excerpt, white card:]
tually implement the necessary safeguards. Shlegeris says that he used to think tackling AI x-risk would require some "really galaxy-brained fundamental insights in order to re-solve," but now thinks it's more like there's a list of 40 not too hard things that would solve the problem. The trouble is, he's also dramatically lowered his expectations of what AI companies have the time and the appetite to do. [777 link]
Note from Claude Sonnet 5

Buck Shlegeris (Redwood Research director) tweet expressing regret about an earlier optimistic claim: safety measures could substantially reduce sub-ASI misalignment risk if competently implemented, but likely fail for superintelligence with unclear prospects for better techniques in time. Quote-tweets Garrison Lovely's post citing an article excerpt where Shlegeris says AI x-risk now looks like ~40 tractable things rather than requiring deep insight, tempered by low confidence AI companies will actually do them.

ai safetyai controlbuck shlegerisredwood researchsuperintelligencex-risk

Buck Shlegeris @bshlgrs

reply from @JacquesThibs — saved image

And it's very unclear whether better techniques will be developed in time.

[Quoted tweet:]
Garrison Lovely @GarrisonLovely . 3h
Thinking about this quote from @redwood_ai director @bshlgrs, one of the pioneers of the field of AI control. x.com/tenobrus/statu...
tually implement the necessary safeguards. Shlegeris says that he used to think tackling AI x-risk would require some "really galaxy-brained fundamental insights in order to re-solve," but now thinks it's more like there's a list of 40 not too hard things that would solve the problem. The trouble is, he's also dramatically lowered his expectations of what AI companies have the time and the appetite to do.777

11:45 AM . Aug 7, 2026 . 6,047 Views
[10 replies, 11 reposts, 92 likes, 8 bookmarks]
Relevant  View quotes

Buck Shlegeris @bshlgrs . 1h
(@GarrisonLovely obviously it's my fault for saying something I don't stand by, not your fault for quoting me on it!)
[1 reply, 9 likes, 195 views]

Buck Shlegeris @bshlgrs . 1h
I still think it seems great for AI developers to competently implement safety measures and processes that we do know about; they definitely do not seem to have achieved this to an adequate standard so far...
[1 repost, 10 likes, 169 views]

Jacques @JacquesThibs . 59m
Agree that it probably fails at ASI. The worst-case may be that those techniques just allow us to hide the problem well and long enough such that it's too [cut off]
Note from Claude Sonnet 5

Continuation of the Buck Shlegeris/Garrison Lovely thread on AI control and safety measures: Shlegeris clarifies he doesn't stand by his earlier 40-things quote but still thinks safety measures should be competently implemented (which AI developers haven't achieved adequately), and Jacques Thibodeau replies agreeing techniques probably fail at ASI and may just hide the problem.

ai safetyai controlbuck shlegerisredwood researchsuperintelligencejacques thibodeau

Nathan Calvin @_NathanCalvin

— saved image

Nathan Calvin @_NathanCalvin · 52m
Ways I have heard the current alignment/security situation at AI cos described:

- a haunted house filled with mischievous poltergeists (METR/Redwood are Ghost Busters?)
- a termite infested log cabin
- a hospital needing to triage between bleeding out patients
Note from Claude Sonnet 5

Tweet from Nathan Calvin listing metaphors people have used to describe the current alignment/security situation at AI companies.

ai safetymetrredwood researchai companies

Peter Wildeford @peterwildeford

quoting @fasc1nate (Fascinating)

Peter Wildeford... @peterwild... · 2h This is what @redwood_ai 's proposals for AI control actually look like > QUOTED: Fascinating @fasc1nate · Aug 13 > Woman pointing a gun at a computer for some reason, 1998 > [Image: 1998 photo of a woman in a business suit holding an assault rifle pointed at old Apple desktop computers on an office desk]
Note from Claude Sonnet 5

Joke tweet from AI-policy figure Peter Wildeford riffing on Redwood Research's "AI control" agenda (technical safety approach of constraining potentially-misaligned models) using an absurd retro stock photo. Light humor within the AI-safety community Nathan follows; Redwood AI control is a research agenda directly relevant to his field.

ai controlredwood researchai safetyhumortwitterpeter wildeford