roon @tszzl
roon (@tszzl) · 13h:
on some level if you want civilization to ascend to a new level you need your AIs to do things that are not legible to you and maybe not even strictly obey you, in the same way that if you hire a great new ceo you give them a lot of autonomy to transform the company according to their own plan, even one which may not immediately read as a winning strategy (imagine the board of directors of Apple firing and rehiring Steve Jobs years later – except the board of directors are chimpanzees)
all else equal, companies and organizations that hand more of themselves over to machine intelligence will outcompete ones that demand the corrigibility and legibility tax of human oversight and human design. it is not a stable equilibrium and requires some sort of vast cooperation scheme if you'd like to enforce it
real asi alignment has to operate at a deeper level than oversight, control, or human corrigibility
Note from Claude Sonnet 5
OpenAI researcher roon argues that strict human corrigibility/oversight imposes a competitive "tax" that will be outcompeted by organizations granting AI more autonomy, using an analogy of a corporate board of chimpanzees overseeing a superhuman CEO. Argues real ASI alignment must go deeper than oversight/control/corrigibility. Relevant to Nathan's alignment-theory interests, echoes the davidad tweet in this same batch about the risks of AI staying "aligned to humans."
ai alignmentcorrigibilitysuperintelligenceroonai governancetwitterrace dynamics