← All topics

singularity tracking

6 captures, most recent first.

Jerry Tworek @MillionInt

reposted by ASM

↻ ASM reposted Jerry Tworek @MillionInt · 2h Recursive self-improvement is here it's just not evenly distributed
Note from Claude Sonnet 5

A short, widely-legible claim from an OpenAI researcher (Jerry Tworek) that recursive self-improvement in AI development is already underway, echoing William Gibson's "the future is here, just not evenly distributed." Directly relevant to Nathan's singularity-r tracking thread.

twitterrecursive self-improvementai timelinesopenaisingularity tracking

Seth Karten @sethkarten

Seth Karten @sethkarten · 14h I only had a 3 month lead over karpathy on auto research. This might change my timelines... I had previously considered this NeurIPS to be the last NeurIPS manageable by human reviewers. I take that back. Now COLM is the last... NeurIPS will be hit with more useful research than it can handle with review demand Either way, GPU demand this year will skyrocket as you are not limited by your management of research agents, but the number of gpus per agent
Note from Claude Sonnet 5

Follow-on commentary in the same "autoresearch" thread (see prior screenshot), a researcher revising AI-driven-research timelines downward and predicting academic peer review will be overwhelmed by AI-generated research volume. Relevant to Nathan's tracking of AI R&D automation and singularity-r indicators.

twitterai r&d automationautonomous research agentspeer reviewtimelinessingularity tracking

snwy @snwy_me

quoting Andrej Karpathy (@karpathy)

snwy @snwy_me · 16h i've been using GPT-5.4 as an autonomous research agent (via Codex) with 24/7 access to an H100 and it has been training/RLing/generating data/repeat a 9B model for the past little while and it is getting crazy fucking good > QUOTED: Andrej Karpathy @karpathy · 16h > I packaged up the "autoresearch" project into a new self-contained minimal repo if people would like to play over the weekend. It's basically nanochat LLM training core stripped down to a single-GPU, one file version of ~630 ... > [Embedded image: chart titled "autoresearch", "Autoresearch Progress: 83 Experiments, 15 Kept Improvements", a step-down line graph of Validation BPB (lower is better) vs Experiment #, showing improvement from ~1.000 to ~0.977 across labeled experiment tweaks (e.g. "raise total batch size", "warmstart LR", "add TF residual", "depth 8 aspect ratio 32"). Caption below: "One day, frontier AI research used to be done by meat computers in between eating, sleeping, having other fun, and synchronizing once in a while using sound wave interconnect in the ritual of 'group meeting'. That era is long gone. Research is now entirely the domain of autonomous swarms of AI agents running across compute cluster megastructures in the skies. The agents claim that we are now in the 10,205th generation of the code base, in any case no one could tell if that's right or wrong as the 'code' is now a self-modifying binary that has grown beyond human comprehension. This repo is the story of how it all began. -@karpathy, March 2026."]
Note from Claude Sonnet 5

Karpathy's "autoresearch" project (an automated LLM-training research loop, satirically captioned as AI agents having fully replaced human researchers) and a user reporting real-world use of GPT-5.4 as an autonomous 24/7 research agent training a 9B model. Directly relevant to Nathan's tracking of AI R&D automation / recursive self-improvement trajectory (cf. Davidson/Houlden singularity-r tracking in memory).

twitterai r&d automationautonomous agentskarpathygpt-5.4recursive self-improvementsingularity tracking

Yuchen Jin @Yuchenj_UW

quote-tweeting METR (@METR_Evals), with embedded quote of roon (@tszzl)

Yuchen Jin @Yuchenj_UW · Jan 29: GPT-5.2 Pro isn't included in the task-length benchmark because it's still thinking and never responded. > QUOTED: METR @METR_Evals · 3h > We're updating the way we measure model time horizons on software tasks (TH 1.0→1.1). The updated methodology incorporates more of the tasks from HCAST, expanding our total from 170 to 228. This produces tighter estimates, especially at longer horizons. > [Chart: "Time Horizon 1.1 and Time Horizon 1.0 show similar 2019-2025 trends, with moderate changes to recent estimates." Y-axis: task length (at 50% success rate), from 4s to 4h+. X-axis: model release date, 2020-2026. Points labeled GPT-2, GPT-3, GPT-3.5, GPT-4, GPT-4o, o1, Claude Opus 4.5 tracing an upward trend line, doubling time 196 days for both TH 1.0 and TH 1.1 methodologies (R²=0.98 and 0.94 respectively).] > 7 comments, 22 reposts, 172 likes, 12K views > [Embedded quote-tweet within: roon @tszzl, 2:54 PM · 1/29/26, 2K views:] > still no 5.2 it's completely over
Note from Claude Sonnet 5

METR's time-horizon benchmark update (TH 1.0→1.1, expanded HCAST task set) tracking AI task-length capability doubling roughly every 196 days, with a joke about GPT-5.2 Pro being too slow to include. Directly relevant to Nathan's tracked "empirical singularity" / METR AI-R&D-automation threads noted in project memory (METR places current automation at 0.25-0.5, doubling time context here is a useful complementary data point).

twittermetrai benchmarkstime horizoncapability scalinggpt-5.2singularity tracking

METR @METR_Evals

METR @METR_Evals "We estimate that, on our tasks, Claude Opus 4.5 has a 50%-time horizon of around 4 hrs 49 mins (95% confidence interval of 1 hr 49 mins to 20 hrs 25 mins). While we're still working through evaluations for other recent models, this is our highest published time horizon to date." [Embedded chart: "Claude Opus 4.5 has a 50%-time horizon of about 4 hrs 49 min (95% CI: 109 to 1225 min)" — log-scale task-length-vs-model-release-date plot from GPT-2 (2019, ~1sec) through GPT-3, GPT-3.5, GPT-4, GPT-4o, o1, o3, GPT-5.1, to Claude Opus 4.5 (~2026, ~4-5 hrs). "Doubling time: 19[?] 2019-01-0[?]"] 6:25 PM · Dec 19, 2025 · 1.2M Views [61 replies, 458 reposts, 1.9K likes, 538 bookmarks] METR @METR_Evals · Dec 19 "We don't think the high upper CI bound reflects Opus's actual capabilities: our current task suite doesn't have enough long tasks to confidently upper bound Opus 4.5's 50%-time horizon. We are working on updating our task suite, and hope to share more details soon." [4 replies, 33 reposts, 448 likes, 76K views] METR @METR_Evals · Dec 19 "Based on our experience interacting with Opus 4.5, the model's performance on specific tasks (including some not in our time horizon suite), and its benchmark performance, we would be surprised..." [cut off]
Note from Claude Sonnet 5

METR's official time-horizon benchmark result for Claude Opus 4.5 — a 50%-task-completion time horizon of ~4 hrs 49 min, their highest published figure to date, continuing the exponential trend on their long-running task-length-vs-capability chart. Directly relevant to the project's "empirical singularity tracking" thread (per project memory: METR places current AI R&D automation at 0.25-0.5x; this Opus 4.5 datapoint extends that trend line into December 2025).

metrtime horizonclaude opus 4.5ai capabilitiesbenchmarkssingularity trackingscaling trends

Rohin Shah @rohinmshah

reposted by David Manheim; reply to @ben_j_todd

↻ David Manheim reposted Rohin Shah @rohinmshah · Dec 24 Replying to @ben_j_todd Both METR and ECI mostly measure things that companies optimize for. 2024 saw the rise of reasoning training for frontier models, which optimizes narrowly for some tasks (whereas pretraining provides more general improvements). So I wouldn't read much into any acceleration.
Note from Claude Sonnet 5

Rohin Shah (DeepMind alignment researcher) pushing back on interpreting METR/ECI capability-benchmark trends as evidence of an acceleration in AI progress, arguing these benchmarks measure exactly what labs already optimize for via reasoning-training, so gains there are less informative than general pretraining improvements would be. Directly relevant to Nathan's tracked "Empirical Singularity Tracking" thread (METR automation estimates, r-value debates) — adds a methodological caveat about benchmark validity that should be logged alongside existing METR/Epoch notes.

twitterai alignmentmetrbenchmarkscapability trendsreasoning trainingsingularity tracking