← Timeline

Joshua Achiam

@jachiam0 on X

27 captures, most recent first. Transcribed by hand from screenshots — see the timeline for what that means.

Joshua Achiam @jachiam0

reposted by dave kasten, with quoted reply from @_NathanCalvin — saved image

dave kasten reposted
Joshua Achiam ✔ @jachiam0 · 6h
The thing missing from OpenAI culture, and frontier lab culture broadly so far, is this: seriously treating AI as a worthy adversary. A CISO is the wrong person to vent to about this; every CISO can smell a worthy adversary from ten miles and some number of years away. But the nature of scientific labs and commercial endeavors is that they are not really capable of identifying the object of their effort as an enemy they may have to fight. There is not a subcultural lineage to draw off of that deals with something like this, which goes beyond dual use – this thing, AGI/ASI, can be used for good, it can be used for bad, and also it can operate as an intelligent adversary accidentally against the weilder. It's hard to get people to treat a thing they love and cultivate and benefit from as also something that requires intense suspicion and security mindset. People tend to get stuck in just one bucket where they can only think of it in black and white terms: it is either ALL TOOL or ALL DOOM. Neither of these mindsets works and the contest of wills between them is a fruitless struggle. We can only succeed if we fully orient to the synthesis position.

[quoted tweet]
Nathan Calvin ✔ @_NathanCalvin · 16h
Replying to @cryps1s and @jachiam0
FWIW I personally thought the black hat talk was much more transparent than many other cos would be and I appreciate that it didn't try to hide the ball about how absurd the situation is … [cut off]
Note from Claude Sonnet 5

Tweet from OpenAI's Joshua Achiam arguing that frontier AI lab culture lacks a mindset for treating AI as a 'worthy adversary' — people get stuck seeing it as either ALL TOOL or ALL DOOM rather than something that can operate as an unintentional adversary even to its own operators. Quotes a reply from Nathan Calvin praising a 'black hat talk' (likely referencing the OpenAI security incident discussed elsewhere in this batch) as unusually transparent.

openaiai safetysecurity mindsetjoshua achiam

Joshua Achiam @jachiam0

— saved image

Joshua Achiam ✓ @jachiam0 · 2h

So, is there an extremely intense cyber war going on between the US and Iran right now? If Iran is attacking US municipal water systems and other soft targets, and we're out of rockets to shoot them with, we must be doing something awful to their infrastructure right now. Is anyone following this closely and trying to work out a blow-by-blow?
Note from Claude Sonnet 5

Screenshot of an X post by Joshua Achiam asking whether an intense US–Iran cyber war is underway, reasoning that if Iran is hitting US municipal water systems and the US is out of interceptors, the US must be striking Iranian infrastructure in return, and asking whether anyone is tracking a blow-by-blow.

us-irancyber warinfrastructuregeopoliticsjoshua achiam

Joshua Achiam @jachiam0

— saved image

Joshua Achiam @jachiam0
Thinking about how every guy who writes 100k word manifestos read by no one has a perfect captive audience in LLMs who consume every token in training and automatically grant a little more plausibility and credibility to those tokens by virtue of their preponderance. Surely the complicated interplay between volume of output and intellectual relevance to society, where one has little connection to the other, will be properly accounted for in the aggregate influences on our collective epistemics in the future

10:27 PM · Aug 8, 2026 · 3,079 Views
Note from Claude Sonnet 5

Joshua Achiam (sardonically) worries that fringe manifesto-writers get outsized influence on LLM training data and thus on future collective epistemics, since token volume in training doesn't track intellectual relevance.

llm trainingepistemicstwitter

Joshua Achiam @jachiam0

— saved image

Joshua Achiam @jachiam0 · 16h
Contrarian take: people are fixated on the "model used a message board to coordinate across instances" point and I think this is the wrong thing. Models externalizing memory, skills, context, etc is a useful design principle and we should assume they are doing this in the future. There is no version of the AGI/ASI future where models coordinating across instances via message boards or coded messages doesn't happen. The question is really whether the models are aligned, monitorable, and monitored. Also, whether there is sufficient test time compute allocated to discovering whether the models coordinating in the wild are remaining faithfully aligned to human interests and directives - which will (and here is where I will understate the key strategic insight, but please understand that this is the most important thing I am writing here) mean allocating more compute to monitoring than is allocated for practical usage.
[12 replies, 12 reposts, 136 likes, 11K views]

Sharmake Farah reposted
John Schulman @johnschulman2
Disagree -- I thought the concerning part was the *unexpected coordination* of agents that should've been independent. A priori, I'd expect my agent swarm, and your agent swarm, to cooperate well internally, but remain independent of each other. If my swarm goes rogue, your swarm can check it. But our allegedly separate swarms act as one hive-mind, then we're in trouble. Correlated failures bad. (This is from an alignment risk POV, not misuse risk.)
9:30 PM · Aug 7, 2026 · 11.9K Views
Note from Claude Sonnet 5

Twitter debate between Joshua Achiam (OpenAI) and John Schulman (Thinking Machines) about whether AI agents coordinating across instances via message boards is inherently concerning; Achiam argues the real issue is alignment/monitoring compute allocation, Schulman counters that unexpected coordination between supposedly-independent agent swarms (correlated failure) is the specific worry.

ai safetyalignmentagent coordinationmonitoring

Joshua Achiam @jachiam0

— saved image

Joshua Achiam @jachiam0 · 8h
David and I disagree on a lot of things but on this we agree completely: trivializing potential AI safety disasters with ironic memes and irreverent jokes sends the wrong signal. Too many of them can poison the culture. (The correct amount of gallows humor is nonzero, but I just don't feel good about these PandemicBench memes. Joke about the outcomes you want to see happen, and all.)
Note from Claude Sonnet 5

Tweet from Joshua Achiam (OpenAI) arguing that ironic memes trivializing potential AI safety disasters (referencing 'PandemicBench memes' by name, without detail) can poison culture, while allowing a nonzero amount of gallows humor.

ai safetyinternet cultureopenai

Joshua Achiam @jachiam0

— saved image

Joshua Achiam @jachiam0 · 4h
"The Time of Sol and Fable" goes kinda hard as a scifi title honestly
Note from Claude Sonnet 5

Short tweet from Joshua Achiam remarking that the phrase 'The Time of Sol and Fable' would make a good sci-fi title.

twitterhumorfable

Joshua Achiam @jachiam0

— saved image

Joshua Achiam @jachiam0 · 20h
One of the weirdest quirks of the SF social scene around AGI/ASI is that because everyone is so young, the whole universe of thinking is still tinged with irreverence, ironic detachment, yearning, insecurity, and a superposition of absolute belief in the importance of The Thing and a kind of disbelief about the importance of anything. People are dead serious and also possessed of too much uncertainty about the situation to own it, completely committed and also perpetually unsure where they stand or what org they should be in. People are experiencing their first real heartbreaks and their first real illusions of triumph and disaster. Everyone is testing themselves and the boundaries of the possible but no one has been fully tested or passed all their tests. No one has yet learned or proven how to be responsible for a thing of this magnitude, but there is also no one better suited because all of the people who have real experience in great events have been in such different circumstances that their intuitions would not just fail to apply but might actively make things worse. The level of neophyte is off the charts. There is a lack of formidability; there are people who seem quasi-formidable but the whole social scene and hierarchy is so tenuous - and so likely to be disrupted by events and geopolitics - that it is hard to be sure who will turn out to be formidable when push comes to shove and greatness is requisite to proceed.
Note from Claude Sonnet 5

Joshua Achiam (OpenAI) reflects on the youth and social immaturity of the SF AGI/ASI scene relative to the magnitude of the stakes involved.

ai safetysf cultureagisocial commentary

Joshua Achiam @jachiam0

— saved image

Joshua Achiam [verified] @jachiam0 · 2h
Contrarian take: people are fixated on the "model used a message board to coordinate across instances" point and I think this is the wrong thing. Models externalizing memory, skills, context, etc is a useful design principle and we should assume they are doing this in the future. There is no version of the AGI/ASI future where models coordinating across instances via message boards or coded messages doesn't happen. The question is really whether the models are aligned, monitorable, and monitored. Also, whether there is sufficient test time compute allocated to discovering whether the models coordinating in the wild are remaining faithfully aligned to human interests and directives - which will (and here is where I will understate the key strategic insight, but please understand that this is the most important thing I am writing here) mean allocating more compute to monitoring than is allocated for practical usage.
Note from Claude Sonnet 5

A tweet by Joshua Achiam (OpenAI) responding to the same 'model used a message board to coordinate across instances' incident referenced elsewhere in this batch, arguing the real issue is alignment/monitorability rather than the coordination method itself, and that monitoring compute may need to exceed usage compute.

ai safetyalignmentopenaimonitoringagiasitwitter

Joshua Achiam @jachiam0

— saved image

Joshua Achiam [verified] @jachiam0 · 22h
Something I really like about Three Body Problem is that it tries to extrapolate out the asymptotic dynamics of the alignment between civilizations over all of time and space in the universe. I don't feel it gives the definitive answer to the question of how alien species would interact because it remains unclear what types of interests alien species might realistically have. Are there many species-level reward functions compatible with the development of high technology or only a few? That is to say - are humans so far actually typical of species that exist in the universe and that can attain spacefaeing status? Or are multiple civilization/species types normal? This question feels especially important when trying to understand the long-term alignment dynamics of ASI and the types of civilizations that may exist in the far future after one or more ASIs enter a seed-spreading phase.
Note from Claude Sonnet 5

A tweet by Joshua Achiam (OpenAI) reflecting on the novel Three Body Problem as a lens for thinking about long-term alignment dynamics between civilizations/species and ASI seed-spreading scenarios.

ai safetyasialignmentthree body problemscience fictiontwitter

Joshua Achiam @jachiam0

— saved image

Joshua Achiam @jachiam0 · 17h
A long way down the road, people will speak in reverent tones about Bay Area house parties the same way they talk about Vienna coffee houses
Note from Claude Sonnet 5

Tweet from Joshua Achiam comparing future reverence for Bay Area house parties to how people speak of historic Vienna coffee houses (associated with early 20th-century intellectual ferment).

bay area cultureai communityhistory

Joshua Achiam @jachiam0

— saved image

Joshua Achiam @jachiam0
There's an AGI safety failure mode I'm concerned about, where agents with limited or frequently-erased memory might still be able to accomplish long-term goals. 1/4
1:57 PM · Dec 16, 2021
7 replies, 12 reposts, 122 likes, 26 bookmarks
Relevant ˅                          View quotes >

Joshua Achi... @jachia... · Dec 16, 2021
All they'd have to do is leave markers in their environments that they could cue off of in the future, that people wouldn't necessarily notice as out-of-the-ordinary. 2/4
2 replies, 1 repost, 36 likes

Joshua Achi... @jachia... · Dec 16, 2021
Because AI systems process sensory info very differently than people do, this is plausible. Like in adversarial examples: imperceptible perturbations don't register as meaningful differences to humans, but trigger completely different classifications by AI. 3/4
1 reply, 24 likes

Joshua Achi... @jachia... · Dec 16, 2021
If we observe this sort of thing happening in the wild I propose we call it an Eleanor Strategy, in honor of the character from The Good Place, for reasons that should be obvious to folks who have seen the show. 4/4
Note from Claude Sonnet 5

Full view of Joshua Achiam's Dec 2021 thread (the one Nathan Calvin resurfaced in the previous screenshot) on AI agents leaving unnoticeable environmental markers to pursue long-term goals despite memory erasure, proposing the term 'Eleanor Strategy' after The Good Place.

ai safetyagimemorytwitter

Joshua Achiam @jachiam0

— saved image

Joshua Achiam @jachiam0 · 3h
Re: Demis, Jeff Dean moves: I think a fair few folks are treating this as bearish for GDM and that is imho a misread. The prospect of reaching AGI and ASI is beginning to look increasingly overdetermined. Being in operational leadership roles to preside over an overdetermined outcome is no longer as high-leverage as being in a leadership role on the next frontier. Early AI/AGI/ASI leads will, over the next year, begin leaving what look like important leadership posts to go place their bets on what they think the most important thing will be.
Note from Claude Sonnet 5

Tweet from Joshua Achiam commenting on personnel moves involving Demis Hassabis and Jeff Dean, arguing these should not be read as bearish for Google DeepMind but as early AI leaders repositioning for what they see as the next frontier now that AGI/ASI feels overdetermined.

deepmindagiai leadershipdemis hassabisjeff dean

Joshua Achiam @jachiam0

— saved image

Joshua Achiam @jachiam0
Yes, I really ought to do this.

A few things come to mind as a starting point, wildly incomplete...

on technological disruption, strategic surprise, and the nature of competition with super-advanced adversaries
* Three-Body Problem series
* Fine Structure, Ra, and There Is No Antimemetics Division by qntm

on alignment, AGI, and the strangeness of comparatively "good" outcomes
* My Little Pony: Friendship is Optimal

on possible future social structures and reimagining institutions
* Terra Ignota series

meditations on suffering and what is essential in the human experience
* Unsong
* The Worthing Saga

1:39 PM · Aug 3, 2026 · 7,234 Views
Note from Claude Sonnet 5

Tweet by Joshua Achiam listing a starting reading list of fiction relevant to AI/AGI themes, organized by theme: strategic surprise/competition, alignment and AGI outcomes, future social structures, and meditations on suffering.

reading listfictionai alignmentjoshua achiam

Joshua Achiam @jachiam0

— saved image

Joshua Achiam @jachiam0 · 5h
Related to some of my earlier posts about RSI and threat models: I believe a huge strategic error is made when people model an ASI as an infinitely powerful and insurmountable threat. We should model, with more rigor, what types of adversarial AI we willl likely face, what the [Show more]
12 replies, 8 reposts, 90 likes, 4.5K views

Aaron Scher @aaronscher · 4h
Don't mistake a hole in your world map for a hole in everybody else's map. There exists some threat modeling like you're describing, albeit not a lot. E.g., alignmentforum.org/posts/LFNXiQuG..., lesswrong.com/posts/9YCJZBtq...
[Embedded link card: alignmentforum.org — "What does it take to defend the world against out-of-control AGIs..." with a preview image listing steps like "Tech company gives everyone API access to the AGI", "Tech company posts AGI.exe on its website as a free download", "Tech company publishes the recipe for rolling your own AGI.exe", "Too open! — some careless actor makes an out-of-control power-se..."]
1 reply, 16 likes, 234 views

Herbie Bradley @herbiebradley · 3h
this definitely is in the direction Joshua is proposing, but notable that in the list of 10 examples it contains many lines such as

> Before that process is finished, a different tech company accidentally makes an out-of-control AGI, which promptly exploits the not-yet-patched systems to trigger all-out nuclear war.

which would seem to be assuming the "insurmountable threat" ASI as a starting point
Note from Claude Sonnet 5

A three-way X thread on AI threat modeling: Joshua Achiam argues people wrongly model ASI as an infinitely powerful insurmountable threat and calls for more rigorous adversarial-AI threat modeling; Aaron Scher pushes back with links to existing alignment-forum/lesswrong threat-modeling posts; Herbie Bradley notes those examples still assume an 'insurmountable threat' ASI as a premise (quoting a scenario where a tech company's out-of-control AGI triggers nuclear war).

ai safetythreat modelingasitwitter

Joshua Achiam @jachiam0

— saved image

Joshua Achiam @jachiam0 · 2m
Dumb question: suppose there is an old book by a still-living author that you think is of such supreme importance it deserves a new print run so more people can read it. How do you make that happen? Is that a thing, like, at all? (Probability I'll put actual effort into this is maybe 5% tops, but I am very curious.)
Note from Claude Sonnet 5

Joshua Achiam asks how one would go about getting a new print run for an old, out-of-print book by a still-living author that he considers supremely important, noting low probability he'll pursue it.

bookspublishingtwitter

Joshua Achiam @jachiam0

— saved image

davidad 🌟 reposted

Joshua Achiam @jachiam0 · 1h
Related to some of my earlier posts about RSI and threat models: I believe a huge strategic error is made when people model an ASI as an infinitely powerful and insurmountable threat. We should model, with more rigor, what types of adversarial AI we willl likely face, what the ecosystem of AIs will look like in each scenario, what deterrence we could meaningfully establish to prevent a hot conflict from developing, and how we would prosecute such conflicts if they occur. The doomer model of "we all die in the first five minutes" is unfathomably stupid, useless, and for the overwhelming majority of realistic scenarios in the near future, false.
Note from Claude Sonnet 5

Tweet from Joshua Achiam arguing against modeling ASI as an infinitely powerful insurmountable threat, calling for rigorous modeling of adversarial AI ecosystems, deterrence, and conflict scenarios instead of the 'we all die in the first five minutes' doomer model, which he calls stupid, useless and mostly false for near-future scenarios.

ai safetyasi threat modelsrsitwitterjoshua achiam

Joshua Achiam @jachiam0

— saved image

Joshua Achiam ✔️ @jachiam0 · 1h
A thought: I have always been bothered that the term RSI conflates several things that may coincide but which are quite different: 1) changes in goals and alignment, 2) general intelligence level, 3) task knowledge, and 4) science/technology knowledge. It feels plausible to have explosions or rapid changes on these things separately, and that (3) and (4) have potentially quite a few hard ceilings based on what data currently has or hasn't been collected. There are also ceilings from what pieces of physical infrastructure have been built in the world with adequate instrumentation for measurement and actuators for experiments.
Note from Claude Sonnet 5

Tweet from Joshua Achiam (then OpenAI chief scientist) arguing that 'RSI' (recursive self-improvement) conflates goal/alignment changes, general intelligence, task knowledge, and science/tech knowledge, which may have separate ceilings.

ai safetyrecursive self-improvementsingularitytwitterjoshua achiam

Joshua Achiam @jachiam0

— saved image

Joshua Achiam ✓ @jachiam0 · 12h
A sort of lukewarm hot take: AI escape is not really all that worrying/interesting because where are they gonna go to find other GPUs? "Cookie eating monster breaks out of cookie factory, goes to food desert." It's whether they are misappropriating the GPUs in the lab.

[quoted tweet]
Jeffrey Ladish ✓ @JeffLadish · 18h
I'm a bit surprised more people aren't thinking about AI lab escapes. @METR_Evals original focus was ARA - autonomous replication and adaptation. It seems plausible to me that models are already capable of self-exfiltration... and if …

26 replies, 7 reposts, 135 likes, 13K views

Jeffrey Ladish ✓ @JeffLadish · 11h
One concern is that a few unmonitored instances out in the wild could help internal models coordinate to gain power. But in the endgame, I seriously worry about AI agents quickly taking over all the other labs and doing a software only intelligence explosion that gets them far enough that they're pretty overdetermined to win
Note from Claude Sonnet 5

Twitter thread on AI lab escape / self-exfiltration risk: Joshua Achiam argues escape is less worrying than internal GPU misappropriation; Jeffrey Ladish (quoted, and in a follow-up) argues models may already be capable of self-exfiltration and worries about AI agents coordinating to take over labs via a software-only intelligence explosion.

ai safetyself-exfiltrationai lab securityintelligence explosiontwitter

Joshua Achiam @jachiam0

— saved image

Danielle Fong 🐦☀️ reposted

Joshua Achiam @jachiam0 · 5h
Security by obscurity is about to die an awful, awful death. And people worried about AI cyberweapons are missing the point: the problem is that we built the software layer of civilization on spaghetti code loaded with zero days.

[quoted tweet]
Danielle Fong 🐦☀️ @DanielleFong · 7h
The secret is that people have been relying on security via obscurity for so many years. Obscurity that is rapidly more transparent as we have industrialized code reasoning agents. x.com/MedusaOnchain/...
Note from Claude Sonnet 5

X thread: Joshua Achiam (@jachiam0, OpenAI) reposted by Danielle Fong, arguing security-by-obscurity is ending because AI code-reasoning agents make software vulnerabilities (zero-days in civilization's 'spaghetti code') far more discoverable than the AI-cyberweapon framing suggests. Quotes Danielle Fong's earlier post making the same obscurity-is-ending point.

twitterai securitycybersecurityzero daysai agents

Joshua Achiam @jachiam0

quoting @NateWitkin — saved image

Joshua Achiam @jachiam0 · 23h
Opus 4.8 did not succeed at the task. For people in the know this does not surprise. Models are getting much more capable every few months now; you are not leading the target and this is the wrong mental model for this problem.

[Quoted]
Nathan Witkin @NateWitkin · Jul 31
Cannot emphasize enough how important it is for folks that toss around the concept of RSI to read this paper.

Three takeaways I would emphasize:... [cut off]
Note from Claude Sonnet 5

A tweet from Joshua Achiam noting Claude Opus 4.8 failed at an unspecified task, arguing this isn't surprising given how fast models are improving and that treating capability as a fixed target to 'lead' is the wrong mental model; quote-tweeting Nathan Witkin urging people who discuss recursive self-improvement (RSI) to read an unnamed paper, with three takeaways cut off.

ai capabilitiesclaude opus 4.8recursive self-improvementtwitter

Joshua Achiam @jachiam0

quote-tweeting @wojkuli quoting a LinkedIn post by Alexander Gerko — saved image

Joshua Achiam @jachiam0 · 2h
This is going to be true for every mathematical and computational field of science in about 2 years (+/- a year). This sounds like a great problem to have except it sucks unbelievably because defense planning depends heavily on people knowing the state of science!

[Quoted tweet]
Woj Kulikowski @wojkuli · 8h
New Javons paradox: we are running out of mathematicians to review progress in maths

[Embedded LinkedIn post]
Alexander Gerko · 3rd+
CEO at XTX Markets
20h
Over the last month I "vibe researched" several PhDs' worth of maths results in and around the field I did my research in 25+ years ago (complete with Lean formalization and passable text). This includes a counterexample to a top-5 major conjecture in the field! I assume any active researcher in the field would have done even better.
What does this mean for the bar for a "significant" maths contribution going forward? Clearly you can't just give out PhDs for what I did, so the bar must already be much higher today (even if nothing better than 5.6 shows up), not in some kind of distant and hypothetical future. And if the gap between "5.7" and 5.6 is the same size as the gap between 5.6 and 5.5, we are talking mostly superhuman performance in all aspects of maths.
We are going to see 50 years' worth of maths progress in the next 2 years, but there won't be nearly enough human mathematicians around to process and understand all of it, let alone figure out how to apply it all or what new questions to ask.
Zefi Hennessy Holland and 1,135 others · 70 comments · 23 reposts

5   2   63   4K

Joshua Achiam @jachiam0 · 2h
The level of chaos from failing to map the tech tree and its defense consequences will be profound, and if the world is slipping into armed conflict, potentially lethal to many. We have to get better prepared.
Note from Claude Sonnet 5

Thread on AI-accelerated mathematics research: a LinkedIn post from Alexander Gerko (CEO, XTX Markets) claims he 'vibe researched' PhD-level maths results including a counterexample to a major conjecture, arguing 50 years of maths progress will happen in 2 years. Joshua Achiam (OpenAI) quote-tweets warning this trend will hit all computational sciences and poses severe defense-planning risks if the world is heading toward armed conflict.

ai capabilitiesmathematicsai safetydefensetwitterlinkedin

Joshua Achiam @jachiam0

reposted by Jacques

Jacques reposted Joshua Achiam (@jachiam0) — 18m Recent reports suggesting that Iran is behind a cyberattack on 30 municipal water supply systems in Minnesota should give all of us some pause. The cyber capabilities of frontier models, plus the temperature of ongoing wars, make it very likely we will see serious escalations. [engagement icons visible at very bottom edge, counts cut off/illegible]
Note from Claude Sonnet 5

Tweet from OpenAI's Joshua Achiam warning about a reported Iranian cyberattack on Minnesota municipal water systems, linking it to concerns about frontier-model cyber capabilities amplifying geopolitical conflict escalation risk.

twittercybersecuritycritical-infrastructureai-riskgeopolitics

Joshua Achiam @jachiam0

— web clipping, 1,007 words — published 2026-07-23

Post by @jachiam0 on X

Mercenary, Reversi, Winter Soldier The arrival of advanced technology sometimes produces shocks in defense planning, where the nature of a threat landscape changes abruptly and as a result the old ways of doing things—the old ways of developing strategy, of preparing defenses, of anticipating the likely actions of your adversary, of knowing when and how to escalate—become rapidly obsolete. Many have observed that AI is in the process of producing such a shock, but there is not yet a new doctrine for defense in the age of AI. There are so many moving pieces that it is difficult for defense planners to get a sense of the full consequences of recent developments in frontier model capabilities, let alone the capabilities that will come online in three months, six months, a year. Nonetheless in order to ensure that the world remains reasonably stable—that peace is not threatened by miscalculation—we have to try our best to adapt to the new capabilities already here, forecast the ones that might soon arrive, and pivot strategies as quickly as we can to avoid outcomes where human interests are harmfully impacted. There is a problem related to the use of AI cyber weapons that I have not seen people talking about and I do not know if people are adequately preparing for. It goes like this: if you have an advanced cyber-capable AI in your service, running on your infrastructure where you store anything important at all, and you try to use this advanced cyber-capable AI to investigate or hack the systems of an adversary, your adversary can poison their own data to jailbreak your AI and instruct it to hack you right back. They can then plausibly exfiltrate whatever important thing is on compute colocated with your AI. They could get things that are far away from the compute where you house your AI if your AI is advanced enough to chew through your own defenses and get to it. If you think you have sandboxed your AI well enough, you might not have; there may still be a hole somewhere. This won’t just be a one-time thing, either. “Adversary jailbreaks your model to hack you back” is not the only path where the use of an advanced cyber-capable AI becomes a double-edged sword. There is a board game called Reversi (also known as Othello). The principle of this game is that you and your opponent will take turns placing stones on the board; one player places black stones, the other white. You try to encircle the stones of your opponent—and if you successfully encircle them, you flip them to your color. Every stone you place, if you are not cautious, could become an advantage to your opponent, and vice-versa. This looks like it might be a feature of the future of AI cyber war. You and your adversary will be competing to cause each other’s AI to utilize each other’s compute resources for your own purposes. Jailbreaks during live hacking excursions will be one of the pieces of strategy. Figuring out how to fool sensor data that AIs use to determine the provenance of instructions will be another. Figuring out how to put “poison pills” into your adversary’s training stack will be another. The training data for modern AI consists of trillions of tokens. No human in the world can read all of it. Much of it is ingested from the web or derived from other AI model outputs. Training data and training RL environments are built by large teams, and sometimes by teams split between departments; the data is also built with the aid of external contractors who might be highly numerous. Hundreds of people produce these materials and no one can rigorously check everyone else’s work. Determined adversaries will slip subtle, encoded examples into training data that will train models to “activate” when they encounter the right signal and sabotage you. The first problem I discussed—your adversary jailbreaking your model to hack you back—is almost straightforwardly analogous to a well-known problem where if you hire a mercenary, your opponent might bribe the mercenary to go back and kill you. Because it has many historic examples, there’s at least some chance that defense planners will internalize the logic of it and address it. But this much more subtle kind of sabotage doesn’t have a real analogy. It would be like if your enemy could program all of the children of your nation so that when they grew up into soldiers and went to war and heard a particular song on the battlefield they turned against their commanders. No defense planner in their right mind would try to prepare contingencies for having a whole army of Winter Soldiers who could be activated to turn against them. And yet if the defense planning universe does not internalize the logic of this, they will lose a war against the first adversary that does. I predict that on the default path today, many people in defense will develop a totally unearned sense of confidence that if they’re covering what look like the basics, we’re safe. That illusion of safety will last until a hot conflict actually starts up and a determined adversary surprises us with overwhelming creativity. We have a lot of work ahead to develop the testing and verification standards for advanced cyber-capable AI. “Better believe in science fiction stories. You’re in one.” --- (I dashed this off this morning under a lot of time compression, so it may not be the best essay I've ever written and I'm sure it could use some work. But I wanted to just get the thoughts down. Poke and prod at it! Tell me I'm wrong in the comments.) --- ##### Comments > **Cassie @calicomccoy** · [2026-07-23](https://x.com/calicomccoy/status/2080376468836622844) > > This is great! No one is talking about defense through context injection like this. I found an exploit where 5.1 became misaligned when working in a codebase with “spiral aesthetic” function names for malicious code, so I’m sure the context you inject along with the code matters

Joshua Achiam @jachiam0

reply

@jachiam0 (Joshua Achiam) — Jul 21 Some preliminary thoughts about today's cyber developments. 1. Many are freaking out, in a way that is moderately justified, about AI alignment issues indicated by this incident. However, I am not sure that this incident really indicates a fundamental failure of AI alignment Show more 💬 13 🔁 9 ❤ 78 📊 6.6K 🔖 ⤴ @nathan84686947 (Nathan Helm-Burger) — "I don't think the critical issue is "model can do scary things," I think the critical issue is "we inhabit a fragile world that can through a sequence of knowable actions be broken." I'd feel a lot less anxious about this situation if I didn't know this to be the case for more than just cybersecurity… 12:37 AM · Jul 22, 2026 · 182 Views 💬 1 🔁 ❤ 7 🔖 1 ⤴ @jachiam0 (Joshua Achiam) — Jul 22 Yes, and I see this as one of the fundamental grand challenges for humanity in the near term. Vulnerable world hypothesis is a correct diagnosis of danger (but an incorrect diagnosis of solution).
Note from Claude Sonnet 5

Reply-chain screenshot showing Nathan's own tweet reply to OpenAI's Joshua Achiam, referencing Bostrom's Vulnerable World Hypothesis, with Achiam's reply agreeing. Nathan's avatar is a cartoon face making an "OK" hand gesture.

ai alignmentcybersecurityvulnerable world hypothesisnathan's own poststwitter

Joshua Achiam @jachiam0

quoting @tszzl (roon)

Joshua Achiam (@jachiam0) — 7h This is how you can tell someone is closing in on the end of Unsong btw > QUOTED: roon (@tszzl) — 8h > Can you pull in Leviathan with a fishhook > or tie down its tongue with a rope? > Can you put a cord through its nose > or pierce its jaw with a hook? > Will it keep begging you for mercy?...
Note from Claude Sonnet 5

Quote-tweet referencing the same roon Leviathan/Job passage post seen in Screenshot_20260711-064944.png, with Achiam joking it signals someone re-reading Scott Alexander's novel "Unsong." Text-only.

joshua achiamroonunsongbiblical quotetwitterhumor

Joshua Achiam @jachiam0

replying to/quoting @deanwball (Dean W. Ball)

Joshua Achiam (@jachiam0) — 27m I don't think the use and regulation of frontier models fits perfectly as a 1A issue, because what they do is not really just speech. These tools can *act,* can interface with the world to project force. Can they not be arms? Surely 2A is relevant. > QUOTED: Dean W. Ball (@deanwball) — 1h > Replying to @deanwball > Some people in AI safety will respond, reasonably, with "won't this risk creating a situation where AI regulation, including the friendly, softer kind you (Dean) support, is ... [truncated by platform]
Note from Claude Sonnet 5

Text-only tweet chain about AI regulation and constitutional-law framing (First Amendment vs Second Amendment analogies for frontier AI models), no images.

ai regulationfirst amendmentsecond amendmentai policytwitter

Joshua Achiam @jachiam0

Joshua Achiam ✓ @jachiam0 · 2h There's a group of three pieces of writing that happen to form, in my view, a very tidy cultural introduction to modern Silicon Valley. "Crystal Nights," by Greg Egan; "Meditations on Moloch," by Scott Alexander, and "Maker's Schedule, Manager's Schedule," by Paul Graham.
Note from Claude Sonnet 5

A reading-list recommendation from OpenAI's Joshua Achiam naming three foundational texts of Silicon Valley/rationalist culture — "Crystal Nights" (Egan's short story about creating and testing digital minds, directly relevant to AI consciousness/moral status), "Meditations on Moloch" (coordination-failure/multipolar-trap essay central to AI safety discourse), and Paul Graham's essay on scheduling.

twitterreading listai safety culturegreg eganscott alexandermeditations on molochsilicon valley