← Timeline

1 capture, most recent first. Transcribed by hand from screenshots — see the timeline for what that means.

@g_leech_

reply by Ben (no treats), @andersonbcdefg

gavin leech (Non-Reasoning) ✔ @g_leech_ improve AI discourse about 5% just by renaming evals accurately Humanity's Last Exam: PubQuizFromHell MATH: RemedialMath FrontierMath: QuarterFrontierMath SWE-Bench: DjangoBench MMLU Virology: NoiseBench Terminal Bench 2: NoiseBench METR HCAST: GreenfieldCodeGigworkBench 4:04 PM · Feb 18, 2026 · 12.1K Views 💬 11 🔁 19 ❤ 377 🔖 56 ⤴ Relevant ⌄ View quotes › Ben (no treats) ✔ @andersonbcdefg · 20h weirdml is accurately named
Note from Claude Sonnet 5

A satirical tweet by Gavin Leech proposing more honest/deflationary renamings for popular AI benchmarks (Humanity's Last Exam, MATH, FrontierMath, SWE-Bench, MMLU Virology, Terminal Bench 2, METR HCAST), critiquing benchmark hype and overclaiming in AI capability discourse. Relevant to Nathan's interest in eval validity and honest capability measurement.

ai benchmarksevalshumortwittercapability measurementgavin leech