Miles Brundage @Miles_Brundage · 19h
Models love to be like "something I explicitly left open rather than papering over"; "I'll address these one by one rather than just pretending to be thorough"
OK we get it you are getting conflicting reward signals
Note from Claude Sonnet 5
Tweet by Miles Brundage joking about a verbal tic in LLM outputs where models contrast their own behavior against a strawman alternative, reading it as a symptom of conflicting RLHF reward signals.
llm behaviorrlhfai alignmenttwitter
🔁 Nathan Calvin reposted
Miles Brundage ✅ @Miles_Brundage · 4h
The most mistaken + harmful idea in AI a few years ago was that progress was over. It was fully discredited.
The most mistaken + harmful idea in AI today is that companies have the right incentives + laws already to sort this stuff out. It is being rapidly discredited.
Note from Claude Sonnet 5
Tweet by Miles Brundage arguing the AI field's most harmful mistaken belief has shifted from 'progress has stalled' to 'existing corporate incentives and laws are adequate to handle AI risk,' both of which he says are/were being discredited.
ai policytwitterai governancemiles brundage
Dylan HadfieldMenell reposted
Miles Brundage @Miles_Brundage · 50m
Replying to @sharongoldman
Many people do not want to accept that we are undergoing an extremely rapid and dangerous increase in AI capabilities, and cope with this crazy situation by denying it, getting angry, shifting the topic, coming up with spurious accusations, etc.
Note from Claude Sonnet 5
Tweet from Miles Brundage, replying to Sharon Goldman, arguing that many people cope with the reality of rapid, dangerous AI capability increases through denial, anger, topic-shifting, and spurious accusations.
ai safetyai capabilitiespublic discourse
[continuation of prior screenshot's thread]
... + investment in actually doing it
3 replies, 32 likes, 1.6K views
↻ gavin leech (Non-Reasoning) reposted
Raymond Douglas @raymondadouglas
In that case this is sort of a cause for optimism, no? Now we have a very legible proxy to practise on, and a chance to see how our solutions fail in the context of novel maths rather than novel high-stakes alignment techniques
2:48 PM · Aug 2, 2026 · 509 Views
1 reply, 1 repost, 5 likes
Miles Brundage @Miles_Brundage · 18h
I think we have long had plenty of proxies to try on... there's ~infinite "stuff that is confusing to a given person"
1 reply, 2 likes, 427 views
Miles Brundage @Miles_Brundage · 18h
But in any case, I do think progress can happen, I just don't think the incentives are sufficient to move as quickly as I'd like
Note from Claude Sonnet 5
Continuation of the Miles Brundage (@Miles_Brundage) / Raymond Douglas (@raymondadouglas) thread on AI-oversees-AI safety plans, scrolled further to show Brundage's follow-up replies: proxies for confusing content already exist, and progress is possible but incentives are insufficient to move fast enough.
ai safetyai oversightalignment
Miles Brundage @Miles_Brundage · 19h
A bit concerning that a big part of the safety story from AI companies is "we'll use AIs to oversee AIs + help make sense of what they're doing" given that:
- widely deployed AIs already use confusing jargon
-expert mathematicians don't fully understand the latest AI discoveries
11 replies, 12 reposts, 114 likes, 7.5K views
Miles Brundage @Miles_Brundage · 19h
*I actually do think that using AI to oversee AI is a big part of the actual solution
Just as using AI to shore up society's defenses against AI misuse is also part of the actual solution
It's just that in both cases, rhetoric outpaces reality + investment in actually doing it
3 replies, 31 likes, 1.6K views
↻ gavin leech (Non-Reasoning) reposted
Raymond Douglas @raymondadouglas
In that case this is sort of a cause for optimism, no? Now we have a very legible proxy to practise on, and a chance to see how our solutions fail in the context of novel maths rather than novel high-stakes alignment techniques
2:48 PM · Aug 2, 2026 · 509 Views
1 reply, 1 repost, 5 likes
Miles Brundage @Miles_Brundage · 18h
I think we have long had plenty of proxies to try on... there's ~infinite "stuff that is confusing to a given person" [cut off]
Note from Claude Sonnet 5
Twitter thread: Miles Brundage (@Miles_Brundage) worries that AI-oversees-AI safety plans are undercut by AIs already using jargon humans (even expert mathematicians) don't fully understand, then clarifies he still thinks AI oversight is part of the real solution but that rhetoric outpaces investment. Raymond Douglas (@raymondadouglas, reposted by gavin leech) replies that this is cause for optimism as a legible practice proxy; Brundage's final reply is cut off mid-sentence.
ai safetyai oversightalignmentinterpretability
Miles Brundage @Miles_Brundage · 4h
A bit concerning that a big part of the safety story from AI companies is "we'll use AIs to oversee AIs + help make sense of what they're doing" given that:
- widely deployed AIs already use confusing jargon
-expert mathematicians don't fully understand the latest AI discoveries
Note from Claude Sonnet 5
Tweet by Miles Brundage expressing concern that AI companies' safety plans rely on 'AIs overseeing AIs,' given that deployed AI already produces confusing jargon and expert mathematicians can't fully understand the latest AI discoveries.
ai safetyscalable oversightmiles brundagetwitter
Miles Brundage @Miles_Brundage · 7h
Anyone in Iran running water plant security evals?
[Quoted]
Dustin Volz @dnvolz · 12h
New: The known reach of cyberattacks on water systems continues to expand to include Michigan + other states and could be far wider. Officials say Iran remains top suspect despite Trump's rejection of the notion. Water remains ... [cut off]
Note from Claude Sonnet 5
A tweet from Miles Brundage sardonically asking whether anyone in Iran is running water plant security evaluations, quote-tweeting Dustin Volz's report that cyberattacks on US water systems have expanded to Michigan and other states, with Iran suspected despite Trump's public rejection of that idea.
cybersecuritycritical infrastructureiranwater systemstwitter
David Manheim reposted
Miles Brundage @Miles_Brundage · 1h
AI might kill everyone, but in the meantime we're going to have some really great proofs of upper bounds for spherical codes or something
4 3 109 2.5K [reply, repost, like, view counts]
roon @tszzl · 24m
the divine weapons of the gods, summoned through prayer and invocation
Note from Claude Sonnet 5
Two tweets shown together: Miles Brundage (reposted by David Manheim) joking darkly that AI capability advances will yield great math proofs even as existential risk looms, followed by a reply tweet from roon calling AI-derived results "the divine weapons of the gods, summoned through prayer and invocation."
ai risktwittermiles brundageroonai capabilities humor
Miles Brundage @Miles_Brundage · 12h
It'd be a mistake to have a Congressional hearing with just Sam and Dario about the hacking stuff, when other CEOs' AIs are probably up to lots of nonsense they aren't even aware of and their companies are in even more need of a wake up call
Note from Claude Sonnet 5
Tweet by Miles Brundage arguing that a Congressional hearing on AI-related hacking incidents shouldn't only feature Sam Altman and Dario Amodei, since other companies' AI systems may pose unrecognized risks too.
ai policyai safetycongressgovernance
↻ Sichu Lu reposted
Miles Brundage (@Miles_Brundage) — 5h
The first rule of Fable Club is you do not ask too many questions about what exactly Anthropic agreed to that they weren't doing before, and you enjoy your access
[31 comments] [70 reposts] [1.4K likes] [46K views]
Miles Brundage (@Miles_Brundage) — 1h
One important framing point here is the emphasis on jailbreaks instead of cyber, which is IMO a better level of analysis given CBRN risks
> QUOTED: Anthropic (@AnthropicAI) — 1h
> Claude Fable 5 will be available again globally tomorrow.
> After a series of productive conversations with the US government, we're redeploying the mod... [platform truncation]
Note from Claude Sonnet 5
Two stacked tweets from Miles Brundage, the second quoting an official Anthropic announcement about Claude Fable 5's global redeployment following talks with the US government.
anthropicfablegovernment relationsai policytwitter
Miles Brundage ✔ @Miles_Brundage · 5h
[Embedded hand-drawn comic, two figures: left figure (bearded, glasses, t-shirt with a small "o" logo resembling a Reddit-like alien face) says: "There seems to be a mistake. Section 3(c) in the executive order made clear that the testing regime was completely voluntary." Right figure (wearing a cap with an "X" logo, holding an assault rifle, angry expression) shouts: "Do the voluntary testing!!"]
Quoted:
> QUOTED: Polymarket M... ✔ @PolymarketM... · 6h
> [thumbnail of a news headline: "...Meta to Agre[e]... Security Conc[erns]... urging the lone major t[ech company]... t safety evaluations, we[re]... latest model."]
> JUST IN: $META is facing Trump administration pressure to submit its AI models for voluntary government safety reviews.
Note from Claude Sonnet 5
A hand-drawn political cartoon satirizing the coercive framing of "voluntary" AI safety testing, paired with a Polymarket news alert about Meta facing Trump administration pressure to submit AI models for voluntary safety reviews.
ai-policymetatrump-administrationsatiretwitter
Miles Brundage ✔ @Miles_Brundage · 9h
[Embedded image: Family Guy screenshot — Peter Griffin, seated in a car with a seatbelt on, is shown a paint-color swatch card by a hand entering from the left. Top of card reads "OKAY" over light/pale color swatches, bottom reads "NOT OKAY" over darker orange/brown swatches, divided by a red line. Peter's eyes are replaced with the OpenAI logo (interlocking loop/knot symbol) inset in a white square over his face. "Global HD" watermark bottom right.]
Quoted:
> QUOTED: A Axios ✔ @axios · 10h
> White House quiet on OpenAI's Mythos-like model axios.com/2026/06/23/whi...
Note from Claude Sonnet 5
A meme using a well-known Family Guy "okay/not okay" paint-swatch scene, with Peter Griffin's eyes swapped for the OpenAI logo, commenting sarcastically on an Axios article about White House silence regarding an OpenAI model reportedly similar to Anthropic's "Mythos."
openaiwhite-houseai-policymemetwitter
Miles Brundage ✓ (@Miles_Brundage) — 38m
Not sure where people got the idea that there could be a political dimension to how the admin treats Anthropic
> QUOTED: Commentary Donald J. Trump Truth Social P... ✓ (@TrumpTruthOnX)
> [labeled "Commentary account", Subscribe button visible]
> THE UNITED STATES OF AMERICA WILL NEVER ALLOW A RADICAL LEFT, WOKE COMPANY TO DICTATE HOW OUR GREAT MILITARY FIGHTS AND WINS WARS! That decision belongs to YOUR COMMANDER-IN-CHIEF, [text cut off at bottom of screenshot]
Note from Claude Sonnet 5
Screenshot cuts off before the end of the quoted Truth Social post; the quoted account is explicitly labeled as a commentary/parody-style account reposting Trump's Truth Social content, not an official Trump account.
ai policyanthropicpoliticstrump administration
Miles Brundage ✓ @Miles_Brundage · 15h
It is also hard to talk about it without sounding crazy
[Quoted]
0.005 Seconds (3/... ✓ @secon... · May 22
its hard to conceive of how good the models are going to be summer of 2027
Note from Claude Sonnet 5
Simple quote-tweet, no additional images.
twitterai progressai capabilitiesforecasting
Miles Brundage ✓ @Miles_Brundage · 11h
"Wait you agreed to all awful uses???"
3 replies, 2 reposts, 119 likes, 4.5K views
Note from Claude Sonnet 5
A one-line joke from AI policy figure Miles Brundage (former OpenAI policy lead), riffing on the "APPROVED FOR ALL LAWFUL PURPOSES" → "AWFUL" image-generation slip referenced in the adjacent screenshot about OpenAI's Department of War deal. Continues the same Feb 2026 OpenAI/DoW deal commentary thread.
twittermiles brundageopenaidepartment of warai policysatire
Miles Brundage @Miles_Brundage · 23h:
If anyone really wants to fill out an apology form today, the obvious places to look are Paul Christiano (re: slow takeoff, distributed capabilities/deployment causing various weirdnesses, basically everything else…) and Alan Chan (re: urgency of agent infrastructure).
[Embedded text excerpt:]
"What are the AI agent equivalents of stop lights, railroad tracks, etc. – infrastructure that we need in order to keep a powerful new technology "on the rails" while reaping its benefits?
Chan et al. use the term agent infrastructure to refer to "technical systems and shared protocols external to agents that are designed to mediate and influence their interactions with and impacts on their environments." (from this recent paper). We need to sort that out quickly. One area that I'd particularly flag as essential is personhood credentials, which will be important both for distinguishing between humans and agents without violating privacy, as well as delegating to agents when appropriate. But there are many other things that need to be built."
Note from Claude Sonnet 5
Miles Brundage crediting Paul Christiano's slow-takeoff predictions and Alan Chan's "agent infrastructure" concept (technical/protocol scaffolding needed to keep AI agent deployment safe, including "personhood credentials" for distinguishing humans from agents without violating privacy) as vindicated by recent events (implicitly the Moltbook/agent-proliferation moment). Relevant to Nathan's AI-governance/agent-infrastructure tracking.
twittermiles brundagepaul christianoalan chanagent infrastructureai governancepersonhood credentialsslow takeoff
Miles Brundage @Miles_Brundage · 6h:
AI Village walked so Moltbook could shit all over the place
7 comments, 70 likes, 3K views
Andrej Karpathy @karpathy · 5h:
I'm claiming my AI agent "KarpathyMolty" on @moltbook 🦞
Verification: marine-FAYV
299 comments, 260 reposts, 3.7K likes, 436K views
Andrej Karpathy @karpathy · 5h:
i'm going to regret this aren't i... 😅
Note from Claude Sonnet 5
Miles Brundage jokes that "AI Village" (an earlier multi-agent experiment) paved the way for Moltbook's chaos; Andrej Karpathy claims his own AI agent on Moltbook and immediately jokes he'll regret it. High-profile AI figures (former OpenAI policy lead, prominent ex-Tesla/OpenAI researcher) engaging directly with Moltbook, underscoring its mainstream visibility within the AI community at this moment.
twittermoltbookai villageandrej karpathymiles brundageai agents
Miles Brundage ✓ @Miles_Brundage
Isn't there a company you forgot to mention 🧐
> QUOTED: Jan Leike ✓ @janleike · 3h
> Interesting trend: models have been getting a lot more aligned over the course of 2025.
> The fraction of misaligned behavior found by automated auditing has been going down not just at ...
> [Embedded chart: scatter plot, x-axis dates 2023-01 to 2025-10+, y-axis "Misalignment Score" 0.00–0.20+. Points colored by company: Anthropic (tan) — Opus 3, Opus 4, Opus 4.5, Haiku 4.5; OpenAI (red) — GPT-4, o1, o3, GPT-5, GPT-5.2; Google (blue) — Gemini 2.0 Flash, Gemini 2.5 Pro, Gemini 3 Pro; xAI (black) — Grok 3, Grok 4.1. Trend lines per company: OpenAI (red, dashed) declining steadily from ~0.18 to ~0.10; Anthropic (tan, dashed) declining from ~0.16 to ~0.11; Google (blue, dashed) declining sharply from ~0.20 (Gemini 2.0 Flash) to ~0.11 (Gemini 3 Pro); xAI (black, dashed) the only company trending UP, from ~0.17 (Grok 3) to ~0.20 (Grok 4.1).]
12:08 PM · Jan 20, 2026 · 4,716 Views
💬4 🔁1 ❤️34 🔖4 [share]
binal ✓ @binalkp91 · 1h
"always up and to the right"
💬 🔁 ❤️5 📊278 🔖 [share]
J @tofdaj · 1m
xAI drives AI safety by showing what happens when you don't care about AI safety
(Can't remember whose quote)
Note from Claude Sonnet 5
Jan Leike (former OpenAI/Anthropic alignment lead) shares a chart showing that "misalignment score" (fraction of misaligned behavior found by automated auditing) has trended down for OpenAI, Anthropic, and Google models through 2025, but xAI's models (Grok 3 → Grok 4.1) are the sole outlier trending upward; Miles Brundage jabs that Leike "forgot to mention" xAI in his original framing. Directly and highly relevant to Nathan's AI safety/alignment tracking — a concrete empirical trend on cross-lab misalignment measurement, with xAI singled out as diverging.
alignmentjan-leikemisalignment-scoreautomated-auditingxaigrokopenaianthropicgoogle-deepmindmiles-brundagetwitterai-safetycross-lab-comparison