← All topics

alignment risk

1 capture, most recent first.

@SharmakeFarah14

replying in thread with @LRudL_ ("Rudolf Laine")

Sharmake Farah @SharmakeFarah14 — 1h A major example is secret loyalties. If we assume that a solution to alignment doesn't require us to have transparency/mech interpretability (as is the case for risk-averse AIs), then companies could safely install secret loyalties in government AIs: forethought.org/research/risk-... > QUOTED: Rudolf Laine ✔️ @LRudL_ — Jun 26 > Replying to @LRudL_ > Of course, if you're an AI lab strategically optimizing for power, this means you go for unprecedented methods to merge with or capture the state
Note from Claude Sonnet 5

Text-only tweet thread about AI transparency/interpretability and the risk of "secret loyalties" embedded in government-deployed AI systems, linking a forethought.org research page.

twitterai interpretabilityai governancestate capturealignment risk