John David Pressman @jd_pressman
— reply from @markopolojarvi (marko.)
John David Pressman ✓ @jd_pressman
The true bitter lesson is if you want the grad student making the model to care about something your critique must come in the form of a benchmark. And the benchmark should have a leaderboard.
> QUOTED (self-reply thread): John David Press... ✓ @jd_pressm... · 4h
> Replying to @jd_pressman
> The guy who makes the model is usually some grad student who cares about almost nothing besides benchmarks. He applies his total genius to marginally improving performance o... Show more
6:34 AM · Apr 9, 2025 · 551 Views
1 reply, 15 likes, 1 bookmark
marko. @markopolojarvi · 4m
What gets measured gets a) gamed and b) fixed.
Note from Claude Sonnet 5
JD Pressman thread arguing that ML research culture is driven by benchmarks/leaderboards rather than qualitative critique — a "bitter lesson" about incentive structures in model-building. Relevant to Nathan's interest in how metrics/incentives shape what gets optimized (echoes Goodhart's-law themes he tracks re: model welfare and alignment).
twittermachine learning culturebenchmarksgoodhart's lawincentivesresearch culture