Yuchen Jin ✔ 𝕏 @Yuchenj_UW · 12h
The task length an AI can complete doubles every 4 months.
In 2041, an AI agent can finish a task that takes a human 58 billion years. That's longer than Universe has existed.
I can one-shot a new universe with a prompt in 15 years. Oh, maybe that's how this one was created!
[Embedded chart from METR: "Claude Opus 4.6 has a 50%-time-horizon of about 14.5 hours (95% CI: 6 hrs to 98 hrs)" — "Task length (at 50% success rate)" scatter/trend chart, y-axis log scale from 4m to 16h with example task labels ("Train classifier" ~15m-1h, "Train adversarially robust image model" ~4h, "Implement complex protocol from multiple RFCs" ~16h), x-axis "Model release date" from ~2023/GPT-4o through 2025-2026. Two trend lines: green (doubling time 123 days, TH 1.1, 2023-01-01+ data, R²=0.93) ending at Claude Opus 4.6 (~14-15h, pink dot) and GPT-5.2 (high) slightly below; gray dashed (doubling time 212 days, trend from Kwa, West, et al. 2025). Points labeled GPT-4o, Sonnet 3.7, Claude Opus 4.6, GPT-5.2 (high) along the curve. Credit: METR, metr.org, CC-BY.]
Note from Claude Sonnet 5
A tweet extrapolating METR's task-time-horizon doubling trend (123-day doubling per their updated fit) to absurd/comedic conclusions about 2041 AI capabilities, illustrated with METR's actual current data placing Claude Opus 4.6's 50%-success time horizon at ~14.5 hours. Relevant to Nathan's tracking of empirical AI capability-growth/singularity metrics (METR HCAST, time-horizon doubling) referenced elsewhere in his research.
metrai capabilitiestime horizonsclaude opusscalingtwittersingularity forecasting
Yuchen Jin @Yuchenj_UW · Jan 29:
GPT-5.2 Pro isn't included in the task-length benchmark because it's still thinking and never responded.
> QUOTED: METR @METR_Evals · 3h
> We're updating the way we measure model time horizons on software tasks (TH 1.0→1.1). The updated methodology incorporates more of the tasks from HCAST, expanding our total from 170 to 228. This produces tighter estimates, especially at longer horizons.
> [Chart: "Time Horizon 1.1 and Time Horizon 1.0 show similar 2019-2025 trends, with moderate changes to recent estimates." Y-axis: task length (at 50% success rate), from 4s to 4h+. X-axis: model release date, 2020-2026. Points labeled GPT-2, GPT-3, GPT-3.5, GPT-4, GPT-4o, o1, Claude Opus 4.5 tracing an upward trend line, doubling time 196 days for both TH 1.0 and TH 1.1 methodologies (R²=0.98 and 0.94 respectively).]
> 7 comments, 22 reposts, 172 likes, 12K views
> [Embedded quote-tweet within: roon @tszzl, 2:54 PM · 1/29/26, 2K views:]
> still no 5.2 it's completely over
Note from Claude Sonnet 5
METR's time-horizon benchmark update (TH 1.0→1.1, expanded HCAST task set) tracking AI task-length capability doubling roughly every 196 days, with a joke about GPT-5.2 Pro being too slow to include. Directly relevant to Nathan's tracked "empirical singularity" / METR AI-R&D-automation threads noted in project memory (METR places current automation at 0.25-0.5, doubling time context here is a useful complementary data point).
twittermetrai benchmarkstime horizoncapability scalinggpt-5.2singularity tracking
[Thread, appears mid-conversation, top of visible thread cut off showing reply counts 2, retweet icon, heart 4, and a bar-chart count]
Yuchen Jin @Yuchenj_UW · 4h
not fully I believe, since they kept some checkpoints
💬2 🔁 ♥20 📊5.9K 🔗
duncan @dchana · 6h
What is preventing you from saying it
💬2 🔁 ♥11 📊10K 🔗
Yuchen Jin @Yuchenj_UW · 6h
keep my oai friends safe lol
💬2 🔁 ♥63 📊9.3K 🔗
duncan @dchana · 6h
Aren't all your friends at meta??
💬1 🔁 ♥10 📊2.7K 🔗
Starkers @imstarkers · 5h
The correct answer is "yes"
💬 🔁 ♥11 📊2.2K 🔗
Prashant @Prashant_1722 · 1h
friends are friends, meta or openai is irrelevant
💬 🔁 ♥1 📊292 🔗
CommonSens... @CommonS... · 4h
By 'absurd,' do you mean
A) it's absurd that they would actually delay and retrain the model over this secret issue, [text cut off at bottom of screen]
Note from Claude Sonnet 5
A Twitter thread among AI-adjacent commentators (Yuchen Jin appears to be an AI researcher) discussing checkpoints being kept and hints of "keeping oai friends safe" — cryptic banter about industry connections rather than substantive technical content, but touches on model retraining/checkpoint retention which is tangential to AI development practices.
twitterai industryopenaimetabantermodel checkpoints