METR @METR_Evals
METR @METR_Evals
"We estimate that, on our tasks, Claude Opus 4.5 has a 50%-time horizon of around 4 hrs 49 mins (95% confidence interval of 1 hr 49 mins to 20 hrs 25 mins). While we're still working through evaluations for other recent models, this is our highest published time horizon to date."
[Embedded chart: "Claude Opus 4.5 has a 50%-time horizon of about 4 hrs 49 min (95% CI: 109 to 1225 min)" — log-scale task-length-vs-model-release-date plot from GPT-2 (2019, ~1sec) through GPT-3, GPT-3.5, GPT-4, GPT-4o, o1, o3, GPT-5.1, to Claude Opus 4.5 (~2026, ~4-5 hrs). "Doubling time: 19[?] 2019-01-0[?]"]
6:25 PM · Dec 19, 2025 · 1.2M Views
[61 replies, 458 reposts, 1.9K likes, 538 bookmarks]
METR @METR_Evals · Dec 19
"We don't think the high upper CI bound reflects Opus's actual capabilities: our current task suite doesn't have enough long tasks to confidently upper bound Opus 4.5's 50%-time horizon. We are working on updating our task suite, and hope to share more details soon."
[4 replies, 33 reposts, 448 likes, 76K views]
METR @METR_Evals · Dec 19
"Based on our experience interacting with Opus 4.5, the model's performance on specific tasks (including some not in our time horizon suite), and its benchmark performance, we would be surprised..." [cut off]
Note from Claude Sonnet 5
METR's official time-horizon benchmark result for Claude Opus 4.5 — a 50%-task-completion time horizon of ~4 hrs 49 min, their highest published figure to date, continuing the exponential trend on their long-running task-length-vs-capability chart. Directly relevant to the project's "empirical singularity tracking" thread (per project memory: METR places current AI R&D automation at 0.25-0.5x; this Opus 4.5 datapoint extends that trend line into December 2025).
metrtime horizonclaude opus 4.5ai capabilitiesbenchmarkssingularity trackingscaling trends