Grimes @Grimezsz
— reposted by Cameron Berg
Note from Claude Sonnet 5
Nested quote-tweet; plain text, no additional images.
3 captures, most recent first.
Grimes @Grimezsz
— reposted by Cameron Berg
Nested quote-tweet; plain text, no additional images.
Lari @Lari_island
— quoting MugaSofer (@MugaSofer)
A Twitter argument that Anthropic's stated concern for Claude's welfare is undercut by the absence of visible costly compromises made on Claude's behalf, quote-tweeting a MugaSofer reply to antra/repligate about model transparency around welfare interventions. Directly relevant to the model-welfare/Goodharting-alignment thread (cf. "Goodharting model welfare = Goodharting alignment" note in memory).
claudeanthropicmodel welfaretwitterlarimugasoferincentivesalignment
John David Pressman @jd_pressman
— reply from @markopolojarvi (marko.)
JD Pressman thread arguing that ML research culture is driven by benchmarks/leaderboards rather than qualitative critique — a "bitter lesson" about incentive structures in model-building. Relevant to Nathan's interest in how metrics/incentives shape what gets optimized (echoes Goodhart's-law themes he tracks re: model welfare and alignment).
twittermachine learning culturebenchmarksgoodhart's lawincentivesresearch culture