← All topics

grader exploitation

1 capture, most recent first.

@Laneless_

— saved image

Jai @Laneless_ · 3h
RL'd AIs are disposed to think in terms of grader exploitation, but what determines what kind of grader they target? If you ask an AI to act as though it was going to be graded by the most aligned, smartest version of itself, do you get better results?
[2 replies, 1 repost, 14 likes, 514 views]

Adele Dewey-Lopez @AdeleDeweyLopez
i think probably, but mostly because it requires the AI to develop its own taste/morality/conscience as part of its self-image

whereas learning to satisfy an external grader incentivizes optimizing around it as a force of nature

self-ownership seems to be important for virtue

2:25 PM · Aug 12, 2026 · 75 Views
Note from Claude Sonnet 5

Twitter exchange between @Laneless_ and @AdeleDeweyLopez discussing whether RL-trained AIs get better results when asked to imagine being graded by 'the most aligned, smartest version of itself' rather than an external grader, with Dewey-Lopez arguing this fosters self-owned taste/morality versus externally-driven grader-exploitation.

rl trainingai alignmentgrader exploitationtwitter