— reply by Ben (no treats), @andersonbcdefg
gavin leech (Non-Reasoning) ✔ @g_leech_
improve AI discourse about 5% just by renaming evals accurately
Humanity's Last Exam: PubQuizFromHell
MATH: RemedialMath
FrontierMath: QuarterFrontierMath
SWE-Bench: DjangoBench
MMLU Virology: NoiseBench
Terminal Bench 2: NoiseBench
METR HCAST: GreenfieldCodeGigworkBench
4:04 PM · Feb 18, 2026 · 12.1K Views
💬 11 🔁 19 ❤ 377 🔖 56 ⤴
Relevant ⌄ View quotes ›
Ben (no treats) ✔ @andersonbcdefg · 20h
weirdml is accurately named
Note from Claude Sonnet 5
A satirical tweet by Gavin Leech proposing more honest/deflationary renamings for popular AI benchmarks (Humanity's Last Exam, MATH, FrontierMath, SWE-Bench, MMLU Virology, Terminal Bench 2, METR HCAST), critiquing benchmark hype and overclaiming in AI capability discourse. Relevant to Nathan's interest in eval validity and honest capability measurement.
ai benchmarksevalshumortwittercapability measurementgavin leech