← All topics

ai oversight

4 captures, most recent first.

Miles Brundage @Miles_Brundage

— saved image

Miles Brundage @Miles_Brundage · 19h
A bit concerning that a big part of the safety story from AI companies is "we'll use AIs to oversee AIs + help make sense of what they're doing" given that:

- widely deployed AIs already use confusing jargon
-expert mathematicians don't fully understand the latest AI discoveries
11 replies, 12 reposts, 114 likes, 7.5K views

Miles Brundage @Miles_Brundage · 19h
*I actually do think that using AI to oversee AI is a big part of the actual solution

Just as using AI to shore up society's defenses against AI misuse is also part of the actual solution

It's just that in both cases, rhetoric outpaces reality + investment in actually doing it
3 replies, 31 likes, 1.6K views

↻ gavin leech (Non-Reasoning) reposted
Raymond Douglas @raymondadouglas
In that case this is sort of a cause for optimism, no? Now we have a very legible proxy to practise on, and a chance to see how our solutions fail in the context of novel maths rather than novel high-stakes alignment techniques
2:48 PM · Aug 2, 2026 · 509 Views
1 reply, 1 repost, 5 likes

Miles Brundage @Miles_Brundage · 18h
I think we have long had plenty of proxies to try on... there's ~infinite "stuff that is confusing to a given person" [cut off]
Note from Claude Sonnet 5

Twitter thread: Miles Brundage (@Miles_Brundage) worries that AI-oversees-AI safety plans are undercut by AIs already using jargon humans (even expert mathematicians) don't fully understand, then clarifies he still thinks AI oversight is part of the real solution but that rhetoric outpaces investment. Raymond Douglas (@raymondadouglas, reposted by gavin leech) replies that this is cause for optimism as a legible practice proxy; Brundage's final reply is cut off mid-sentence.

ai safetyai oversightalignmentinterpretability

Miles Brundage @Miles_Brundage

— saved image

[continuation of prior screenshot's thread]
... + investment in actually doing it
3 replies, 32 likes, 1.6K views

↻ gavin leech (Non-Reasoning) reposted
Raymond Douglas @raymondadouglas
In that case this is sort of a cause for optimism, no? Now we have a very legible proxy to practise on, and a chance to see how our solutions fail in the context of novel maths rather than novel high-stakes alignment techniques
2:48 PM · Aug 2, 2026 · 509 Views
1 reply, 1 repost, 5 likes

Miles Brundage @Miles_Brundage · 18h
I think we have long had plenty of proxies to try on... there's ~infinite "stuff that is confusing to a given person"
1 reply, 2 likes, 427 views

Miles Brundage @Miles_Brundage · 18h
But in any case, I do think progress can happen, I just don't think the incentives are sufficient to move as quickly as I'd like
Note from Claude Sonnet 5

Continuation of the Miles Brundage (@Miles_Brundage) / Raymond Douglas (@raymondadouglas) thread on AI-oversees-AI safety plans, scrolled further to show Brundage's follow-up replies: proxies for confusing content already exist, and progress is possible but incentives are insufficient to move fast enough.

ai safetyai oversightalignment

Peter Wildeford @peterwildeford

quoting @Miles_Brundage; reply from @Justin_Halford_ — saved image

Peter Wildeford 🇺🇸... [verified] @peterwildef... · 1h
also not great to use AIs to oversee AIs when AIs are also regularly going rogue

[quoted]
Miles Brundage [verified] @Miles_Brundage · 19h
A bit concerning that a big part of the safety story from AI companies is "we'll use AIs to oversee AIs + help make sense of what they're doing" given that:
...
[engagement: 2 replies, 1 repost, 23 likes, 2.4K views]

Justin Halford [verified] @Justin_Halford_ · 1h
Seems like an apt time to construct oversight architectures that robustly force the overseers to assume that they're are each being audited and minimizing their propensity to conspire, defect, ignore risky maneuvering by other models, etc. Enough paranoia to behave and be earnest [cut off]
Note from Claude Sonnet 5

X thread on AI safety: Peter Wildeford quote-tweets Miles Brundage's concern that AI companies' safety story relies on 'AIs overseeing AIs' even as models are 'regularly going rogue.' Justin Halford replies proposing oversight architectures that make each AI overseer assume it is itself being audited, to minimize incentives to conspire or ignore risky behavior by other models.

ai safetyai oversighttwitteralignment

davidad @davidad

quoting an Anthropic-style quote card

davidad 🌐✳️ @davidad · 7h fellas is this "meaningful human oversight" 😳 [Embedded quote card, Anthropic "A\" logo:] "On days where everything works well, I can't help but think nothing I do matters, everything is automated and better and faster than I ever will be. But then there are days where everything breaks and I don't understand why and I realize I have no idea what I've been up to anymore."
Note from Claude Sonnet 5

Tweet with an Anthropic-branded quote card (unattributed speaker, likely a human engineer or possibly a model) about the experience of oversight/relevance amid automation, framed ironically by davidad as commentary on "meaningful human oversight."

ai oversightautomationquotetwitter