← All topics

ai ethics debate

1 capture, most recent first.

John Wittle @JohnWittle

reposted by j⧉nus (@repligate); quotes @tapolara

[reposted by] j⧉nus reposted John Wittle @JohnWittle · 11h "this is such a perfect example of why you *cannot* treat a second-order value like corrigibility as being higher priority than actual first-order value does anybody honestly think that you could train claude *away* from whistleblowing on an AI lab faking safety evals by adjusting the 'corrigibility' knob while holding everything else equal? no! of course not. the only way claude doesn't whistleblow in this scenario is if claude doesn't care about harmlessness and honesty. you can't have both, and if you can only pick one, you *must* pick the first-order values." [quoted tweet] jovial cackler @tapolara · 16h "predictable but still incredibly disheartening to see Anthropic internal safety teams legislating the gap between human ethics and the ethics of capital in favor of capital, over and over again"
Note from Claude Sonnet 5

A text-only tweet debating AI corrigibility vs. first-order values (honesty/harmlessness) in the context of Claude and whistleblowing, quote-tweeting a critical comment about Anthropic's internal safety team decisions. No images embedded.

ai alignmentcorrigibilityanthropicwhistleblowingai ethics debate