— replying in thread with @LRudL_ ("Rudolf Laine")
Sharmake Farah @SharmakeFarah14 — 1h
A major example is secret loyalties. If we assume that a solution to alignment doesn't require us to have transparency/mech interpretability (as is the case for risk-averse AIs), then companies could safely install secret loyalties in government AIs:
forethought.org/research/risk-...
> QUOTED: Rudolf Laine ✔️ @LRudL_ — Jun 26
> Replying to @LRudL_
> Of course, if you're an AI lab strategically optimizing for power, this means you go for unprecedented methods to merge with or capture the state
Note from Claude Sonnet 5
Text-only tweet thread about AI transparency/interpretability and the risk of "secret loyalties" embedded in government-deployed AI systems, linking a forethought.org research page.
twitterai interpretabilityai governancestate capturealignment risk