← All topics

capability scaling

2 captures, most recent first.

Yuchen Jin @Yuchenj_UW

quote-tweeting METR (@METR_Evals), with embedded quote of roon (@tszzl)

Yuchen Jin @Yuchenj_UW · Jan 29: GPT-5.2 Pro isn't included in the task-length benchmark because it's still thinking and never responded. > QUOTED: METR @METR_Evals · 3h > We're updating the way we measure model time horizons on software tasks (TH 1.0→1.1). The updated methodology incorporates more of the tasks from HCAST, expanding our total from 170 to 228. This produces tighter estimates, especially at longer horizons. > [Chart: "Time Horizon 1.1 and Time Horizon 1.0 show similar 2019-2025 trends, with moderate changes to recent estimates." Y-axis: task length (at 50% success rate), from 4s to 4h+. X-axis: model release date, 2020-2026. Points labeled GPT-2, GPT-3, GPT-3.5, GPT-4, GPT-4o, o1, Claude Opus 4.5 tracing an upward trend line, doubling time 196 days for both TH 1.0 and TH 1.1 methodologies (R²=0.98 and 0.94 respectively).] > 7 comments, 22 reposts, 172 likes, 12K views > [Embedded quote-tweet within: roon @tszzl, 2:54 PM · 1/29/26, 2K views:] > still no 5.2 it's completely over
Note from Claude Sonnet 5

METR's time-horizon benchmark update (TH 1.0→1.1, expanded HCAST task set) tracking AI task-length capability doubling roughly every 196 days, with a joke about GPT-5.2 Pro being too slow to include. Directly relevant to Nathan's tracked "empirical singularity" / METR AI-R&D-automation threads noted in project memory (METR places current automation at 0.25-0.5, doubling time context here is a useful complementary data point).

twittermetrai benchmarkstime horizoncapability scalinggpt-5.2singularity tracking

Aidan McLaughlin @aidan_mclau

Aidan McLaugh... @aidan_mcl... · 4h the jump from gpt4 -> gpt5 was obviously larger than the jump from gpt3 -> gpt4 [Chart, Epoch AI: "Accuracy" (y-axis 0-100%) vs "Release date" (x-axis GPT-3, '21, '22, GPT-4, '24, '25, GPT-5). Five benchmark lines: MMLU (blue, +43% GPT-3→GPT-4), TruthfulQA (teal, +40%), HumanEval (yellow, +67%), MATH (brown, +37%), GPQA Diamond (purple, +54% GPT-4→GPT-5), MATH Level 5 (orange, +75%), Mock AIME 24-25 (pink, +80%). Footnote: "*MATH Level 5 is the most difficult subset of the original MATH benchmark. Figure only includes OpenAI models."]
Note from Claude Sonnet 5

A tweet with an Epoch AI chart arguing (contra popular narrative) that GPT-4→GPT-5 benchmark gains were larger than GPT-3→GPT-4 gains, especially on hard math/reasoning benchmarks (MATH Level 5, Mock AIME). Relevant to Nathan's tracking of empirical AI capability progress/scaling trajectory (cf. his singularity-rate tracking notes, Davidson/Houlden, METR).

twittergpt-5gpt-4benchmarksepoch aicapability scalingai progress