← All topics

human oversight

2 captures, most recent first.

Tenobrus @tenobrus

— saved image

[Link card, continued from prior screenshot]
en.wikipedia.org
Andon (manufacturing) - Wikipedia
2 replies, 1 repost, 46 likes, 909 views

Tenobrus @tenobrus · 3h
yeah i really think we need to start including infra and rewards for things like "asking for clarification", "surfacing confusion", "pinging your human manager with an update", etc. all things we very much want weak AGI to actually do
1 reply, 4 reposts, 77 likes, 783 views

swisscheese @swisscheese4299 · 28m
🎭 Commentary account
This is a good thought.
7 views

Kromem @kromem2dot0 · 2h
Which is why it's important to find a balance in adding availability of human outreach to the infra.

In my own work deployments I have a suggestion box for the agents and reporting pathways for issues that might arise. If I didn't add these they wouldn't assume it was an option.
Note from Claude Sonnet 5

Continuation of the Yudkowsky Twitter thread on AI agents never being trained to consult humans. Tenobrus's reply is now shown complete; swisscheese (a 'commentary account') briefly agrees; Kromem describes adding a suggestion box and reporting pathways to their own agent deployments so agents know human outreach is an option, since otherwise they wouldn't assume it.

ai alignmenttwitterclarifying questionsagent deploymenthuman oversight

Marius Hobbhahn @MariusHobbhahn

@MariusHobbha... (Marius Hobbha...) — 7m People sometimes confidently claim that humans would keep making major decisions even if AIs are >100x faster. Imagine you could only chat with your boss on one day per year! a) it would be very clear to everyone that this is not workable b) you'd just make decisions around your boss and disempower them in order to get anything done. I expect the situation with AIs will look comparable, especially if they are rewarded based on their outcomes.
Note from Claude Sonnet 5

Single tweet, dark mode, no images.

ai speedhuman oversightai safetydisempowerment