← All topics

rob miles

3 captures, most recent first.

Rob Miles @robertskmiles

— saved image

Rob Miles @robertskmiles · 23h
Being very abusive to chatbots is interesting, because it's plausibly a morally neutral act, but it's done exclusively by bad people
Note from Claude Sonnet 5

A tweet by AI safety communicator Rob Miles observing that abusing chatbots is plausibly morally neutral yet done exclusively by bad people.

ai moral statuschatbot abusetwitterrob miles

Eliezer Yudkowsky @allTheYud

— saved image

Eliezer Yudkowsky @allTheYud
My current thought: AIs are never RLed on working with real humans; that would be expensive. AIs are never RLed on a task where they can consult a human and get help. So AIs solemnly debate among themselves, and 0 in 10,000 consider "talk to a human" as an option.

[Quoted tweet]
Eliezer Yudkowsky @allTheYud · Aug 8
A confusion: Thousands of GPTs debated among themselves which crimes ought or ought not be committed. Zero defected / whistleblew / told a human.
...

4:17 PM · Aug 11, 2026 · 19.5K Views
31 replies, 19 reposts, 416 likes, 72 bookmarks

Relevant  |  View quotes

Rob Miles @robertskmiles · 3h
I think training should include an 'Andon Cord' tool, to allow the agent to flag problems with the task etc. It may still help even if it doesn't always go to a real human during training, as long as the incentive structure is right
[Link card: en.wikipedia.org — Andon (manufacturing) - Wikipedia]
2 replies, 1 repost, 46 likes, 909 views

Tenobrus @tenobrus · 3h
yeah i really think we need to start including infra and rewards for things like "asking for clarification", "surfacing confusion", "pinging your human manager [cut off]
Note from Claude Sonnet 5

Twitter thread led by Eliezer Yudkowsky arguing AIs are never RL-trained on tasks involving real human consultation, so they never learn to ask for help, illustrated by his earlier claim that zero of thousands of GPTs debating crime ethics chose to whistleblow to a human. Rob Miles replies proposing an 'Andon Cord' tool (referencing the manufacturing andon cord concept) letting agents flag problems during training. Tenobrus agrees more infra/reward should exist for clarification-seeking and surfacing confusion to a human manager, reply cut off.

ai alignmenteliezer yudkowskyrob milesrlhftwitterclarifying questionsandon cord

Rob Miles @robertskmiles

Rob Miles @robertskmiles · 5m: "New AI risk discovered: When speech recognition results in your personal assistant not "marking those to-dos as done" but "mocking those to-dos as dumb""
Note from Claude Sonnet 5

A joke tweet by AI safety communicator Rob Miles about a speech-recognition mishap, unrelated to substantive safety content — light humor in Nathan's feed.

twitterhumorai assistantsrob miles