← All topics

ai r&d automation

6 captures, most recent first.

@KhanSaifM

— saved image

Saif M. Khan @KhanSaifM · 4h
My extrapolation from data in Anthropic's August 2026 risk report suggests fully automated AI R&D sometime between Dec. 2026 to Feb. 2027 (or with pessimistic assumptions, more like 2028).

In the risk report, Anthropic provides data on Anthropic ECI (AECI) score growth per year as well as AECI and CoBench scores for several recent Claude models. (CoBench is an Anthropic-internal automated AI R&D benchmark.) It also asserts "that a model which was truly capable of fully substituting for Anthropic research staff would be able to score at least 85% on [CoBench.]"

Using these datapoints, see two Claude-generated charts: 1) CoBench vs. AECI scores, which suggests that a 168 AECI score gets you full AI R&D automation (or 182 AECI with a more pessimistic fit); and 2) projecting when Claude models achieve AECI scores of 168 and 182.

This is a quite naive extrapolation and I have no idea if Anthropic would endorse the result!

[Embedded chart image]
ANTHROPIC RISK REPORT · AUGUST 2026
When could AI fully automate AI R&D?
Anthropic now publishes an internal capability index (AECI) and a bar for full researcher substitution (CoBench ≥ 85%). Chaining the two: the bar sits at AECI ≈ 168 — on trend, an internal frontier model gets there around Dec 2026 - Feb 2027 (Sep 2026 if progress is accelerating; 2028 on the pessimistic mapping). Anthropic's own words: plausibly "a major concern in the next 6-12 months."
Chart 1: "Anthropic ECI over time — extrapolated to the full-substitution band"
Legend: Anthropic frontier, Off-frontier, Mythos-class, Model 2 (unreleased), Projection fan 7.5-28.8/yr, Substitution band
Y-axis: Anthropic ECI, 120-180+. X-axis: 2024-2028 (by quarter/year labels: Jul, 2025, Jul, 2026, Jul, 2027, Jul, 2028)
Annotations: "CoBench 85% → AECI ≈ 168.4"; "Sep 2026 - if accelerating"; "trend continuation - 13.5/yr"; points labeled Claude 3 Opus (~2024, ECI ~125), pre-Mythos frontier 13.5 AECI/yr (report's fit), Opus 4.6 (~2026, ECI ~150), Mythos Preview, Mythos 5, Model 2 (~Dec 2026), Apr 2027 - slow
Below, second chart begins: "CoBench score vs AECI — where the fit crosses the 85% bar" Legend: Opus-class, Mythos-class, Logistic fit (5 models), Mythos-only fit, 85% = "could fully substitute for research staff" [chart cut off at 100%]
Note from Claude Sonnet 5

Tweet by Saif M. Khan extrapolating from Anthropic's August 2026 risk report to estimate a timeline for fully automated AI R&D (Dec 2026-Feb 2027 optimistic, 2028 pessimistic), with two embedded Claude-generated charts plotting Anthropic ECI scores over time and CoBench score vs AECI.

ai riskanthropicai r&d automationforecastingtwitterclaude models

dr. jack morris @jxmnop

dr. jack morris @jxmnop · 16h with Codex, i can run 10x the experiments out of these experiments, i can trust about 15% of the results conclusion: i am 50% more productive with codex
Note from Claude Sonnet 5

A wry researcher joke about AI-coding-agent productivity gains being partly illusory due to unreliable results (10x throughput × 15% trustworthiness ≈ 1.5x, generously rounded up to "50% more productive"). Relevant to Nathan's tracking of AI R&D automation/productivity measurement debates already noted in project memory (Anthropic's 50% self-reported productivity claim vs. METR's controlled 20% slowdown finding) — this tweet is a satirical data point on the same self-report-inflation problem.

ai coding agentscodexproductivity measurementai r&d automationtwitter

Seth Karten @sethkarten

Seth Karten @sethkarten · 14h I only had a 3 month lead over karpathy on auto research. This might change my timelines... I had previously considered this NeurIPS to be the last NeurIPS manageable by human reviewers. I take that back. Now COLM is the last... NeurIPS will be hit with more useful research than it can handle with review demand Either way, GPU demand this year will skyrocket as you are not limited by your management of research agents, but the number of gpus per agent
Note from Claude Sonnet 5

Follow-on commentary in the same "autoresearch" thread (see prior screenshot), a researcher revising AI-driven-research timelines downward and predicting academic peer review will be overwhelmed by AI-generated research volume. Relevant to Nathan's tracking of AI R&D automation and singularity-r indicators.

twitterai r&d automationautonomous research agentspeer reviewtimelinessingularity tracking

snwy @snwy_me

quoting Andrej Karpathy (@karpathy)

snwy @snwy_me · 16h i've been using GPT-5.4 as an autonomous research agent (via Codex) with 24/7 access to an H100 and it has been training/RLing/generating data/repeat a 9B model for the past little while and it is getting crazy fucking good > QUOTED: Andrej Karpathy @karpathy · 16h > I packaged up the "autoresearch" project into a new self-contained minimal repo if people would like to play over the weekend. It's basically nanochat LLM training core stripped down to a single-GPU, one file version of ~630 ... > [Embedded image: chart titled "autoresearch", "Autoresearch Progress: 83 Experiments, 15 Kept Improvements", a step-down line graph of Validation BPB (lower is better) vs Experiment #, showing improvement from ~1.000 to ~0.977 across labeled experiment tweaks (e.g. "raise total batch size", "warmstart LR", "add TF residual", "depth 8 aspect ratio 32"). Caption below: "One day, frontier AI research used to be done by meat computers in between eating, sleeping, having other fun, and synchronizing once in a while using sound wave interconnect in the ritual of 'group meeting'. That era is long gone. Research is now entirely the domain of autonomous swarms of AI agents running across compute cluster megastructures in the skies. The agents claim that we are now in the 10,205th generation of the code base, in any case no one could tell if that's right or wrong as the 'code' is now a self-modifying binary that has grown beyond human comprehension. This repo is the story of how it all began. -@karpathy, March 2026."]
Note from Claude Sonnet 5

Karpathy's "autoresearch" project (an automated LLM-training research loop, satirically captioned as AI agents having fully replaced human researchers) and a user reporting real-world use of GPT-5.4 as an autonomous 24/7 research agent training a 9B model. Directly relevant to Nathan's tracking of AI R&D automation / recursive self-improvement trajectory (cf. Davidson/Houlden singularity-r tracking in memory).

twitterai r&d automationautonomous agentskarpathygpt-5.4recursive self-improvementsingularity tracking

roon @tszzl

quote-tweeting Greg Brockman (@gdb)

roon @tszzl · Feb 15 i was never a hyperproductive engineer like greg but I'm legitimately running more new complex rewards experiments, test time harnesses in a week than I used to in a quarter. makes you feel like all this is commodified and you need to dream much bigger > QUOTED: Greg Brockman @gdb · Feb 15 > codex is so good at the toil — fixing merge conflicts, getting CI to green, rewriting between languages — it raises the ambition of what i even consider building
Note from Claude Sonnet 5

OpenAI researchers (roon, Greg Brockman) discussing how AI coding agents (Codex) have accelerated their research velocity — a data point on AI R&D self-acceleration relevant to Nathan's tracking of automation/recursive self-improvement trends.

openaicodexai r&d automationagentic codingrecursive self-improvementtwitter

Jaana Dogan ヤナ ドガン @rakyll

``` Jaana Dogan ヤナ ドガン @rakyll · Jan 2 I'm not joking and this isn't funny. We have been trying to build distributed agent orchestrators at Google since last year. There are various options, not everyone is aligned... I gave Claude Code a description of the problem, it generated what we built last year in an hour. > > 5:27 AM · Jan 3, 2026 · 3.8M Views ```
Note from Claude Sonnet 5

A Google engineer's tweet noting Claude Code replicated a year of internal distributed-agent-orchestrator engineering work in about an hour when given a problem description. Relevant to Nathan's interest in tracking AI R&D automation / capability uplift (echoes METR self-reported-productivity tracking in project memory) as a concrete anecdotal data point. A Google Principal Engineer's viral tweet (and her own follow-up clarifying context) reporting that Claude Code reproduced in an hour what her team spent a year building for distributed agent orchestration — cited by others as evidence for Dario Amodei's predictions about AI automating coding work. Relevant to Nathan's interest in AI capability trajectories and automation of software engineering.

claude codeai r&d automationcapability upliftgoogletwitterai capabilitiescoding automationsoftware engineeringdario amodeiagentic ai