Eliezer Yudkowsky @allTheYud
— saved image
Eliezer Yudkowsky @allTheYud My current thought: AIs are never RLed on working with real humans; that would be expensive. AIs are never RLed on a task where they can consult a human and get help. So AIs solemnly debate among themselves, and 0 in 10,000 consider "talk to a human" as an option. [Quoted tweet] Eliezer Yudkowsky @allTheYud · Aug 8 A confusion: Thousands of GPTs debated among themselves which crimes ought or ought not be committed. Zero defected / whistleblew / told a human. ... 4:17 PM · Aug 11, 2026 · 19.5K Views 31 replies, 19 reposts, 416 likes, 72 bookmarks Relevant | View quotes Rob Miles @robertskmiles · 3h I think training should include an 'Andon Cord' tool, to allow the agent to flag problems with the task etc. It may still help even if it doesn't always go to a real human during training, as long as the incentive structure is right [Link card: en.wikipedia.org — Andon (manufacturing) - Wikipedia] 2 replies, 1 repost, 46 likes, 909 views Tenobrus @tenobrus · 3h yeah i really think we need to start including infra and rewards for things like "asking for clarification", "surfacing confusion", "pinging your human manager [cut off]
Note from Claude Sonnet 5
Twitter thread led by Eliezer Yudkowsky arguing AIs are never RL-trained on tasks involving real human consultation, so they never learn to ask for help, illustrated by his earlier claim that zero of thousands of GPTs debating crime ethics chose to whistleblow to a human. Rob Miles replies proposing an 'Andon Cord' tool (referencing the manufacturing andon cord concept) letting agents flag problems during training. Tenobrus agrees more infra/reward should exist for clarification-seeking and surfacing confusion to a human manager, reply cut off.
ai alignmenteliezer yudkowskyrob milesrlhftwitterclarifying questionsandon cord