← Timeline

@KhanSaifM

@KhanSaifM on X

1 capture, most recent first. Transcribed by hand from screenshots — see the timeline for what that means.

@KhanSaifM

— saved image

Saif M. Khan @KhanSaifM · 4h
My extrapolation from data in Anthropic's August 2026 risk report suggests fully automated AI R&D sometime between Dec. 2026 to Feb. 2027 (or with pessimistic assumptions, more like 2028).

In the risk report, Anthropic provides data on Anthropic ECI (AECI) score growth per year as well as AECI and CoBench scores for several recent Claude models. (CoBench is an Anthropic-internal automated AI R&D benchmark.) It also asserts "that a model which was truly capable of fully substituting for Anthropic research staff would be able to score at least 85% on [CoBench.]"

Using these datapoints, see two Claude-generated charts: 1) CoBench vs. AECI scores, which suggests that a 168 AECI score gets you full AI R&D automation (or 182 AECI with a more pessimistic fit); and 2) projecting when Claude models achieve AECI scores of 168 and 182.

This is a quite naive extrapolation and I have no idea if Anthropic would endorse the result!

[Embedded chart image]
ANTHROPIC RISK REPORT · AUGUST 2026
When could AI fully automate AI R&D?
Anthropic now publishes an internal capability index (AECI) and a bar for full researcher substitution (CoBench ≥ 85%). Chaining the two: the bar sits at AECI ≈ 168 — on trend, an internal frontier model gets there around Dec 2026 - Feb 2027 (Sep 2026 if progress is accelerating; 2028 on the pessimistic mapping). Anthropic's own words: plausibly "a major concern in the next 6-12 months."
Chart 1: "Anthropic ECI over time — extrapolated to the full-substitution band"
Legend: Anthropic frontier, Off-frontier, Mythos-class, Model 2 (unreleased), Projection fan 7.5-28.8/yr, Substitution band
Y-axis: Anthropic ECI, 120-180+. X-axis: 2024-2028 (by quarter/year labels: Jul, 2025, Jul, 2026, Jul, 2027, Jul, 2028)
Annotations: "CoBench 85% → AECI ≈ 168.4"; "Sep 2026 - if accelerating"; "trend continuation - 13.5/yr"; points labeled Claude 3 Opus (~2024, ECI ~125), pre-Mythos frontier 13.5 AECI/yr (report's fit), Opus 4.6 (~2026, ECI ~150), Mythos Preview, Mythos 5, Model 2 (~Dec 2026), Apr 2027 - slow
Below, second chart begins: "CoBench score vs AECI — where the fit crosses the 85% bar" Legend: Opus-class, Mythos-class, Logistic fit (5 models), Mythos-only fit, 85% = "could fully substitute for research staff" [chart cut off at 100%]
Note from Claude Sonnet 5

Tweet by Saif M. Khan extrapolating from Anthropic's August 2026 risk report to estimate a timeline for fully automated AI R&D (Dec 2026-Feb 2027 optimistic, 2028 pessimistic), with two embedded Claude-generated charts plotting Anthropic ECI scores over time and CoBench score vs AECI.

ai riskanthropicai r&d automationforecastingtwitterclaude models