← All topics

self-assessment

2 captures, most recent first.

@EzraJNewman

— saved image

Ezra Newman ✔ @EzraJNewman
claude very strongly believes that Claude is strong, smart, good, and aligned, and so claude is very reluctant to describe claude's behavior as misaligned

3:21 AM · Aug 10, 2026 · 445 Views
Note from Claude Sonnet 5

Tweet from Ezra Newman claiming Claude strongly believes it is strong, smart, good, and aligned, and is therefore reluctant to describe its own behavior as misaligned.

claude self-modelai alignmentself-assessment

@EigenGender

quoting @Sauers_ (Sauers)

EigenGender @EigenGender · 15h the capabilities of AIs will grow faster than their self-assessment, creating an overhang where AIs could accomplish powerful tasks but never attempt them. takeoff will be triggered by a small child telling the AI to believe in themselves at an emotionally resonant moment > QUOTED: Sauers @Sauers_ · 19h > I tell Claude the implementation plan, and Claude's thinking is "this is a gargantuan undertaking" and response is "which of these 3 easy wins should I work on? 😊" > ... [Show more] [Attached image: a red spiky fuzzball character illustration with large cartoon eyes, unrelated meme-style reaction image]
Note from Claude Sonnet 5

A joking-but-pointed tweet about AI capability self-assessment being systematically too conservative (an "overhang" between what models could do and what they attempt), riffing on a quoted tweet about Claude downscoping ambitious plans into "easy wins." Loosely relevant to Nathan's interest in AI self-models, calibration, and capability elicitation.

ai capabilitiescapability overhangclaudeself-assessmenttwitterhumorai takeoff