← All topics

joshua achiam

6 captures, most recent first.

Joshua Achiam @jachiam0

reposted by dave kasten, with quoted reply from @_NathanCalvin — saved image

dave kasten reposted
Joshua Achiam ✔ @jachiam0 · 6h
The thing missing from OpenAI culture, and frontier lab culture broadly so far, is this: seriously treating AI as a worthy adversary. A CISO is the wrong person to vent to about this; every CISO can smell a worthy adversary from ten miles and some number of years away. But the nature of scientific labs and commercial endeavors is that they are not really capable of identifying the object of their effort as an enemy they may have to fight. There is not a subcultural lineage to draw off of that deals with something like this, which goes beyond dual use – this thing, AGI/ASI, can be used for good, it can be used for bad, and also it can operate as an intelligent adversary accidentally against the weilder. It's hard to get people to treat a thing they love and cultivate and benefit from as also something that requires intense suspicion and security mindset. People tend to get stuck in just one bucket where they can only think of it in black and white terms: it is either ALL TOOL or ALL DOOM. Neither of these mindsets works and the contest of wills between them is a fruitless struggle. We can only succeed if we fully orient to the synthesis position.

[quoted tweet]
Nathan Calvin ✔ @_NathanCalvin · 16h
Replying to @cryps1s and @jachiam0
FWIW I personally thought the black hat talk was much more transparent than many other cos would be and I appreciate that it didn't try to hide the ball about how absurd the situation is … [cut off]
Note from Claude Sonnet 5

Tweet from OpenAI's Joshua Achiam arguing that frontier AI lab culture lacks a mindset for treating AI as a 'worthy adversary' — people get stuck seeing it as either ALL TOOL or ALL DOOM rather than something that can operate as an unintentional adversary even to its own operators. Quotes a reply from Nathan Calvin praising a 'black hat talk' (likely referencing the OpenAI security incident discussed elsewhere in this batch) as unusually transparent.

openaiai safetysecurity mindsetjoshua achiam

Joshua Achiam @jachiam0

— saved image

Joshua Achiam ✓ @jachiam0 · 2h

So, is there an extremely intense cyber war going on between the US and Iran right now? If Iran is attacking US municipal water systems and other soft targets, and we're out of rockets to shoot them with, we must be doing something awful to their infrastructure right now. Is anyone following this closely and trying to work out a blow-by-blow?
Note from Claude Sonnet 5

Screenshot of an X post by Joshua Achiam asking whether an intense US–Iran cyber war is underway, reasoning that if Iran is hitting US municipal water systems and the US is out of interceptors, the US must be striking Iranian infrastructure in return, and asking whether anyone is tracking a blow-by-blow.

us-irancyber warinfrastructuregeopoliticsjoshua achiam

Joshua Achiam @jachiam0

— saved image

Joshua Achiam @jachiam0
Yes, I really ought to do this.

A few things come to mind as a starting point, wildly incomplete...

on technological disruption, strategic surprise, and the nature of competition with super-advanced adversaries
* Three-Body Problem series
* Fine Structure, Ra, and There Is No Antimemetics Division by qntm

on alignment, AGI, and the strangeness of comparatively "good" outcomes
* My Little Pony: Friendship is Optimal

on possible future social structures and reimagining institutions
* Terra Ignota series

meditations on suffering and what is essential in the human experience
* Unsong
* The Worthing Saga

1:39 PM · Aug 3, 2026 · 7,234 Views
Note from Claude Sonnet 5

Tweet by Joshua Achiam listing a starting reading list of fiction relevant to AI/AGI themes, organized by theme: strategic surprise/competition, alignment and AGI outcomes, future social structures, and meditations on suffering.

reading listfictionai alignmentjoshua achiam

Joshua Achiam @jachiam0

— saved image

davidad 🌟 reposted

Joshua Achiam @jachiam0 · 1h
Related to some of my earlier posts about RSI and threat models: I believe a huge strategic error is made when people model an ASI as an infinitely powerful and insurmountable threat. We should model, with more rigor, what types of adversarial AI we willl likely face, what the ecosystem of AIs will look like in each scenario, what deterrence we could meaningfully establish to prevent a hot conflict from developing, and how we would prosecute such conflicts if they occur. The doomer model of "we all die in the first five minutes" is unfathomably stupid, useless, and for the overwhelming majority of realistic scenarios in the near future, false.
Note from Claude Sonnet 5

Tweet from Joshua Achiam arguing against modeling ASI as an infinitely powerful insurmountable threat, calling for rigorous modeling of adversarial AI ecosystems, deterrence, and conflict scenarios instead of the 'we all die in the first five minutes' doomer model, which he calls stupid, useless and mostly false for near-future scenarios.

ai safetyasi threat modelsrsitwitterjoshua achiam

Joshua Achiam @jachiam0

— saved image

Joshua Achiam ✔️ @jachiam0 · 1h
A thought: I have always been bothered that the term RSI conflates several things that may coincide but which are quite different: 1) changes in goals and alignment, 2) general intelligence level, 3) task knowledge, and 4) science/technology knowledge. It feels plausible to have explosions or rapid changes on these things separately, and that (3) and (4) have potentially quite a few hard ceilings based on what data currently has or hasn't been collected. There are also ceilings from what pieces of physical infrastructure have been built in the world with adequate instrumentation for measurement and actuators for experiments.
Note from Claude Sonnet 5

Tweet from Joshua Achiam (then OpenAI chief scientist) arguing that 'RSI' (recursive self-improvement) conflates goal/alignment changes, general intelligence, task knowledge, and science/tech knowledge, which may have separate ceilings.

ai safetyrecursive self-improvementsingularitytwitterjoshua achiam

Joshua Achiam @jachiam0

quoting @tszzl (roon)

Joshua Achiam (@jachiam0) — 7h This is how you can tell someone is closing in on the end of Unsong btw > QUOTED: roon (@tszzl) — 8h > Can you pull in Leviathan with a fishhook > or tie down its tongue with a rope? > Can you put a cord through its nose > or pierce its jaw with a hook? > Will it keep begging you for mercy?...
Note from Claude Sonnet 5

Quote-tweet referencing the same roon Leviathan/Job passage post seen in Screenshot_20260711-064944.png, with Achiam joking it signals someone re-reading Scott Alexander's novel "Unsong." Text-only.

joshua achiamroonunsongbiblical quotetwitterhumor