← Timeline

1 capture, most recent first. Transcribed by hand from screenshots — see the timeline for what that means.

Rohin Shah @rohinmshah

reposted by David Manheim; reply to @ben_j_todd

↻ David Manheim reposted Rohin Shah @rohinmshah · Dec 24 Replying to @ben_j_todd Both METR and ECI mostly measure things that companies optimize for. 2024 saw the rise of reasoning training for frontier models, which optimizes narrowly for some tasks (whereas pretraining provides more general improvements). So I wouldn't read much into any acceleration.
Note from Claude Sonnet 5

Rohin Shah (DeepMind alignment researcher) pushing back on interpreting METR/ECI capability-benchmark trends as evidence of an acceleration in AI progress, arguing these benchmarks measure exactly what labs already optimize for via reasoning-training, so gains there are less informative than general pretraining improvements would be. Directly relevant to Nathan's tracked "Empirical Singularity Tracking" thread (METR automation estimates, r-value debates) — adds a methodological caveat about benchmark validity that should be logged alongside existing METR/Epoch notes.

twitterai alignmentmetrbenchmarkscapability trendsreasoning trainingsingularity tracking