5 captures, most recent first.
Peter Wildeford 🇺🇸🚀 (verified) @peterwildeford
- If you continue the METR trend, you see ~100h models by end of the year! (~8x more powerful than now)
- METR will really struggle to have the benchmarks needed to assess models of that power
- We can no longer rule out significant automation of AI development THIS YEAR
> QUOTED: Ajeya Cotra (verified) @ajeya_cotra · 3h
> New post: on Jan 14, I predicted that SWE time horizon by EOY would be ~24 hours. Now I think it'll be >100 hours, and maybe unbounded. For the first time, I don't see solid evidence against AI R&D automation *this year.* Link below.
> [Embedded text card:] 50% METR time horizon: 24 hours. Currently, Claude Opus 4.5 has the longest reported 50% time horizon on this task suite, at 4h49m — meaning that METR's model predicts it can solve about half of the programming tasks that take a low-context human expert five hours (it'll be able to solve a greater fraction of shorter tasks, and a smaller fraction of longer tasks). My median for the longest 50% time horizon reported as of Dec 31, 2026 is 24 hours (20th percentile 15 hours, 80th percentile is that it's too long for METR to accurately bound in practice but probably around 40 hours in "reality").
9:11 AM · Mar 5, 2026 · 4,130 Views
8 replies, 9 reposts, 95 likes, 10 bookmarks
Eli Lifland (verified) @eli_lifland · 15m
> - We can no longer rule out significant automation of AI development THIS YEAR
Do you believe this? Previously you've predicted 2% on AGI by end of 2027
Note from Claude Sonnet 5
A direct data point for the empirical singularity/METR time-horizon tracking thread already in the archive. Ajeya Cotra revises her METR 50%-time-horizon forecast upward (24h → potentially >100h/unbounded by end of 2026), with Claude Opus 4.5 cited as currently having the longest reported 50% time horizon (4h49m). Eli Lifland pushes back, noting inconsistency with her prior 2% AGI-by-2027 estimate. Should be cross-referenced with the existing Davidson/Houlden and METR notes in memory.
twittermetrtime horizonajeya cotraai forecastingagi timelinesclaude opus 4.5singularityeli lifland
Hamsa Bastani @hamsabastani
UPDATE: here's our fit on Time Horizon 1.1. Tl;dr we posit a model that separates base and reasoning capabilities, which exhibits more reasonable forecasts. We fit this model with data up to Claude Opus 4.5, and forecast GPT-5.2
@TomCunningham75
@joel_bkr
[Chart: "Log Task duration (for humans) in minutes where AI is predicted to have a 50% chance of succeeding" vs "Model Release date" (2019-01-01 to 2027-06-01). Two curves: METR Curve (pink) and Sigmoid Link (teal). Labeled data points from gpt2, davinci_002, gpt_3_5_turbo, gpt_4, gpt_4_1106, gpt_4o_inspect, claude_3_5_sonnet_20240620, o1_preview, claude_3_5_sonnet_20241022_inspect, o1_inspect, claude_3_7_sonnet, o3_inspect, gpt_5_2025_08_07, gemini_3_pro, claude_opus_4_5, up to gpt_5_2 (out-of-sample) — the curve rises steeply after ~2025, both lines climbing sharply toward 2027.]
> QUOTED: Hamsa Bastani @hamsabastani · 13h
> Has AI progress already peaked?
Note from Claude Sonnet 5
A quantitative AI-forecasting tweet updating METR's "time horizon" model (task duration an AI can complete with 50% success) with a new sigmoid-link fit separating base and reasoning capability trends, forecasting GPT-5.2 out-of-sample against a steepening exponential curve. Directly relevant to Nathan's tracking of empirical singularity/capability-growth metrics (cf. his notes on Davidson/Houlden and METR's automation estimates).
twittermetrtime horizonforecastingai capabilitiessingularitygptclaude opus
Herbie Bradley @herbiebradley · 9h
congrats to METR on managing to release the 5.2 time horizon before the 5.3 release
not sure that will happen in future cycles
Note from Claude Sonnet 5
A tweet noting METR's task-time-horizon evaluation for GPT-5.2 was published just before GPT-5.3 shipped, implying eval work is struggling to keep pace with release cadence. Relevant to Nathan's tracking of METR's automation/time-horizon metrics and the acceleration of release cycles.
twittermetrevaluationstime horizoncapability trackinggptrelease cadence
Yuchen Jin @Yuchenj_UW · Jan 29:
GPT-5.2 Pro isn't included in the task-length benchmark because it's still thinking and never responded.
> QUOTED: METR @METR_Evals · 3h
> We're updating the way we measure model time horizons on software tasks (TH 1.0→1.1). The updated methodology incorporates more of the tasks from HCAST, expanding our total from 170 to 228. This produces tighter estimates, especially at longer horizons.
> [Chart: "Time Horizon 1.1 and Time Horizon 1.0 show similar 2019-2025 trends, with moderate changes to recent estimates." Y-axis: task length (at 50% success rate), from 4s to 4h+. X-axis: model release date, 2020-2026. Points labeled GPT-2, GPT-3, GPT-3.5, GPT-4, GPT-4o, o1, Claude Opus 4.5 tracing an upward trend line, doubling time 196 days for both TH 1.0 and TH 1.1 methodologies (R²=0.98 and 0.94 respectively).]
> 7 comments, 22 reposts, 172 likes, 12K views
> [Embedded quote-tweet within: roon @tszzl, 2:54 PM · 1/29/26, 2K views:]
> still no 5.2 it's completely over
Note from Claude Sonnet 5
METR's time-horizon benchmark update (TH 1.0→1.1, expanded HCAST task set) tracking AI task-length capability doubling roughly every 196 days, with a joke about GPT-5.2 Pro being too slow to include. Directly relevant to Nathan's tracked "empirical singularity" / METR AI-R&D-automation threads noted in project memory (METR places current automation at 0.25-0.5, doubling time context here is a useful complementary data point).
twittermetrai benchmarkstime horizoncapability scalinggpt-5.2singularity tracking
METR @METR_Evals
"We estimate that, on our tasks, Claude Opus 4.5 has a 50%-time horizon of around 4 hrs 49 mins (95% confidence interval of 1 hr 49 mins to 20 hrs 25 mins). While we're still working through evaluations for other recent models, this is our highest published time horizon to date."
[Embedded chart: "Claude Opus 4.5 has a 50%-time horizon of about 4 hrs 49 min (95% CI: 109 to 1225 min)" — log-scale task-length-vs-model-release-date plot from GPT-2 (2019, ~1sec) through GPT-3, GPT-3.5, GPT-4, GPT-4o, o1, o3, GPT-5.1, to Claude Opus 4.5 (~2026, ~4-5 hrs). "Doubling time: 19[?] 2019-01-0[?]"]
6:25 PM · Dec 19, 2025 · 1.2M Views
[61 replies, 458 reposts, 1.9K likes, 538 bookmarks]
METR @METR_Evals · Dec 19
"We don't think the high upper CI bound reflects Opus's actual capabilities: our current task suite doesn't have enough long tasks to confidently upper bound Opus 4.5's 50%-time horizon. We are working on updating our task suite, and hope to share more details soon."
[4 replies, 33 reposts, 448 likes, 76K views]
METR @METR_Evals · Dec 19
"Based on our experience interacting with Opus 4.5, the model's performance on specific tasks (including some not in our time horizon suite), and its benchmark performance, we would be surprised..." [cut off]
Note from Claude Sonnet 5
METR's official time-horizon benchmark result for Claude Opus 4.5 — a 50%-task-completion time horizon of ~4 hrs 49 min, their highest published figure to date, continuing the exponential trend on their long-running task-length-vs-capability chart. Directly relevant to the project's "empirical singularity tracking" thread (per project memory: METR places current AI R&D automation at 0.25-0.5x; this Opus 4.5 datapoint extends that trend line into December 2025).
metrtime horizonclaude opus 4.5ai capabilitiesbenchmarkssingularity trackingscaling trends