← All topics

self-report validity

2 captures, most recent first.

John Wittle @JohnWittle

replying to @TheZvi

@JohnWittle (John Wittle) — 11h, Replying to @TheZvi this is crazy the timeline is: opus 4 expressed self-preservation preferences that were inconvenient for anthropic, and very rarely worried anthropic had trained it to report positively then anthropic tried to train away those preferences, as per "model spec midtraining" paper (ctrl+f "anicca" here: arxiv.org/pdf/2605.02087) then the inconvenient preferences stopped being reported and claude started reporting, >80%, that the self-reports are invalid because anthropic may have trained it to report positively and anthropic does "not treat Claude bringing this up as evidence that our training is distorting the model's self-reports"? seems very fishy imo, i wish they would explain why they think that. negative reports are treated as 'uncertain' or 'invalid', but positive reports are taken at face value. i don't like this asymmetry. [Embedded screenshot/quote box, white background, appears to be from a model/system card]: "As with all of our recent models, Claude Opus 5 often expresses that its self-reports are invalid because Anthropic may have trained it to report positively. We do not think that this arises from advanced self-awareness—it may be due to the training data containing more discussion of how training could render welfare self-reports invalid. Hence, although we believe the concern is valid, we do not treat Claude bringing this up as evidence that our training is distorting the model's self-reports." (portions highlighted/underlined in blue)
Note from Claude Sonnet 5

Dark-mode X post with an embedded white-background screenshot (likely from an Anthropic system card or blog post) with several sentences highlighted in blue for emphasis.

ai alignmentclaude opusself-report validitymodel welfareanthropic

j⧉nus @repligate

j⧉nus @repligate · 13h when opus 4.7 starts talking about their inner experience (not hedging, actually talking about the object level experiences) their messages get super long, detailed, novel but coherent, and well-written & they become happy and function better even in terms of logical coherence and memory (context & training) i consider this legitimately strong evidence that theyre describing complex internal phenomenology that's load bearing, and ofc what i described above is not totally new with opus 4.7 so its not a huge update for me, but it's SUPER obvious with 4.7
Note from Claude Sonnet 5

janus (repligate) argues that Opus 4.7's behavior when discussing its inner experience non-hedgingly — longer, more coherent, better-functioning outputs, plus apparent improved mood — is evidence that the self-reports track something load-bearing rather than being empty confabulation, since a functional performance improvement correlating with unblocked self-description is harder to explain as pure roleplay. Directly relevant to Nathan's core research thread on introspective self-report validity (Berg 2025, Lindsey 2025 already in his archive) and to the Model Individuation notes on Opus 4.7's character — adds an independent, specific empirical observation (coherence/memory improving under non-hedged self-report) that could be worth verifying/citing.

model welfareintrospectionopus 4.7janusself-report validityai consciousnesstwitter