7 captures, most recent first.
Epoch AI @EpochAIResearch · 53m
We've updated the MirrorCode leaderboard with results for Claude Fable 5 and GPT-5.6 Sol.
Claude Fable 5 leads with a 64% solve rate, followed by GPT-5.6 Sol at 20%.
[Chart: "AI models can autonomously complete some large software projects" — Overall score on MirrorCode (ML, +Private, 2L). Even when they fail to reimplement targets, AIs typically make substantial progress, passing 90% or more of tests. Bar chart, solve@100% rate (per-target mean), whiskers ±1 SE: Claude Fable 5 64%, GPT-5.6 Sol 20%, GPT-5.4 16%, GPT-5.5 10%. EPOCH AI | CC-BY, epoch.ai]
13 replies, 32 reposts, 213 likes, 8.3K views
Epoch AI @EpochAIResearch · 53m
MirrorCode tests whether AI agents can reimplement software projects from scratch. A solve requires passing 100% of visible and hidden tests.
Claude Fable is the first model we evaluated to solve the C preprocessor and Pkl tasks in at least one run.
[Chart: "The hardest MirrorCode targets remain unsolved" — Per-target solve rates for MirrorCode (ML, +Private, 2L), for full reimplementation (100% of tests passed). Table by model x target:
Claude Fable 5 (avg 64% ±10pp): tssql 100, private_M 100, texmacros 100, wren_cli 100, bib2json 100, nonogrid 83, brotild 83, gotree 0, sed 67, mailauth 83, giac_subset 0, cprepro 83, pkl 25, private_L 33, ruff 0.
GPT-5.6 Sol (avg 20% ±9pp): tssql 100, private_M 50, texmacros 0, wren_cli 33, bib2json 100, nonogrid 0, brotild 17, gotree 0, sed 0, mailauth 0, giac_subset 0, cprepro 0, pkl 0, private_L 0, ruff 0.
GPT-5.4 (avg 16% ±8pp): tssql 100, private_M 50, texmacros 0, wren_cli 0, bib2json 50, nonogrid 0, brotild 33, gotree 0, sed 0, mailauth 0, giac_subset 0, cprepro 0, pkl 0, private_L 0, ruff 0.
GPT-5.5 (avg 10% ±6pp): tssql 83, private_M 0, texmacros 0, wren_cli 0, bib2json 33, nonogrid 0, brotild 33, gotree 0, sed 0, mailauth 0, giac_subset 0, cprepro 0, pkl 0, private_L 0, ruff 0.
Footnote: Each model was run three times in each of the two implementation languages for a target (six planned attempts per target). Attempts with infrastructure errors and no valid final evaluation are excluded; scored sample counts are shown in tooltips. Overall scores average the 15 per-target solve rates, weighting each target equally. ± values show 1 SE across target programs. EPOCH AI | CC-BY, epoch.ai]
Note from Claude Sonnet 5
Two-tweet Epoch AI thread announcing updated MirrorCode leaderboard results: Claude Fable 5 leads with a 64% solve rate on full software-project reimplementation, well ahead of GPT-5.6 Sol (20%), GPT-5.4 (16%), and GPT-5.5 (10%), including per-target breakdown tables.
ai benchmarkingclaude fablesoftware engineeringepoch ai
[repost icon] Sichu Lu reposted
@timfduffy (Tim Duffy) — 50m
If you use the middle of each provided range as the mean for that bucket, total contributed hours are ~1.5x as high as they were a year ago. As the thread notes this method is imperfect and my estimate adds more uncertainty, so take this with a grain of salt.
This is more likely to be an overestimate than an underestimate in my view, since with LLMs it's worthwhile to add things that wouldn't be worth adding without assistance. So the time to create estimates probably rise more than value created.
[Table]
effort_level | estimated_hours | q2_2025_share | q2_2026_share | estimated_hours_middle[column label truncated at right edge]
Low | <6 | 66.3 | 50.9 | 3[possibly truncated]
Medium | 6-12 | 18.8 | 24 | 9[possibly truncated]
High | 12-24 | 12.9 | 16.9 | 18[possibly truncated]
Very high | 24-48 | 2 | 7.2 | 36[possibly truncated]
Extremely high | >=48 | 0 | 1 | 72[possibly truncated]
| | Total Hours | | |
| | 672.3 | 1004.1 | |
| | Speedup Factor | | |
| | 1.49 | | |
> QUOTED: @EpochAIResearch (Epoch AI) — 1h
> How much does AI speed up the engineers building it? We analyzed contributions to OpenAI's public Codex repository to gather evidence. ... [truncated by platform]
> [Image: bar chart thumbnail, not legible at this resolution]
Note from Claude Sonnet 5
A tweet analyzing Epoch AI's research on AI-driven engineer productivity using OpenAI's public Codex repository; Tim Duffy recomputes a "speedup factor" of ~1.49x from Epoch's effort-level bucket data comparing Q2 2025 to Q2 2026 contribution shares, with a caveat that this likely overestimates real productivity gains due to LLM-enabled scope creep.
ai productivitytwitterepoch aiai forecastingsoftware engineeringdata analysis
@magnushambleton (magnus) — 14h
I am so so tired of all the moaning about Dario and regulation, please can it stop
Epoch AI currently estimates open source models are 4 months behind the frontier.
So if you can't see any merit in being even slightly careful or thoughtful during the short span of time when we are creating a new species, just wait 4 months and stop moaning.
Note from Claude Sonnet 5
Standalone tweet referencing Dario (Amodei, implied) and Epoch AI's open-source lag estimate.
ai regulationdario amodeiopen source aiepoch ai
Thomas Woodsi... @Thomas_Wo... · 15h
A lot of graphs with this shape lately
Epoch AI @EpochAIResearch · 18h
AI companies say their models are getting better at finding software vulnerabilities. Is that bearing out in public data?
Introducing our Cyber Vulnerabilities explorer, ...
[Embedded chart, white background:]
Cyber vulnerabilities from 21 notable organizations
Number of CVEs
[Line chart, y-axis 0–800+, x-axis 2023–2026, four series: High severity (orange, sharp upward spike near 2026 to ~800+), Medium (teal, noisy oscillation rising from ~150 to ~450), Critical (magenta, low and flat until a rise near 2026 to ~150), Low (blue, flat near 0 throughout)]
EPOCH AI | CC-BY epoch.ai
Note from Claude Sonnet 5
Line chart showing a sharp late-2025/2026 upward spike specifically in "High severity" CVE counts among 21 organizations, framed by the poster as an example of a recurring exponential-looking AI-capability chart shape.
ai capabilitiescybersecurityepoch aicveresearch
Jaime Sevilla @Jsevillamol · 8h
Anthropic still on track to be the first to a gigawatt datacenter online.
> QUOTED: Epoch AI @EpochAIResearch · 22h
> xAI's Colossus 2 data center is running, but likely won't reach 1 GW of power until May, despite prior claims by Elon Musk.
>
> Our updated analysis shows the facility lacks t...
> [Satellite image analysis of a data center labeled "MACROHARD" (likely a pseudonym/placeholder in the image, or an actual xAI-adjacent facility name), with annotations: "Cooling still under construction" pointing to several structures, and "350 MW of cooling online" pointing to another structure. Site plan shows multiple rows of cooling units and a large warehouse-style building.]
Note from Claude Sonnet 5
A tweet comparing datacenter buildout progress between AI labs (Anthropic vs. xAI's Colossus 2), citing Epoch AI's satellite-imagery analysis. Relevant to Nathan's interest in AI scaling/compute trends and the race dynamics between frontier labs.
ai scalingdatacenterscomputexaianthropicepoch aitwitter
Aidan McLaugh... @aidan_mcl... · 4h
the jump from gpt4 -> gpt5 was obviously larger than the jump from gpt3 -> gpt4
[Chart, Epoch AI: "Accuracy" (y-axis 0-100%) vs "Release date" (x-axis GPT-3, '21, '22, GPT-4, '24, '25, GPT-5). Five benchmark lines: MMLU (blue, +43% GPT-3→GPT-4), TruthfulQA (teal, +40%), HumanEval (yellow, +67%), MATH (brown, +37%), GPQA Diamond (purple, +54% GPT-4→GPT-5), MATH Level 5 (orange, +75%), Mock AIME 24-25 (pink, +80%). Footnote: "*MATH Level 5 is the most difficult subset of the original MATH benchmark. Figure only includes OpenAI models."]
Note from Claude Sonnet 5
A tweet with an Epoch AI chart arguing (contra popular narrative) that GPT-4→GPT-5 benchmark gains were larger than GPT-3→GPT-4 gains, especially on hard math/reasoning benchmarks (MATH Level 5, Mock AIME). Relevant to Nathan's tracking of empirical AI capability progress/scaling trajectory (cf. his singularity-rate tracking notes, Davidson/Houlden, METR).
twittergpt-5gpt-4benchmarksepoch aicapability scalingai progress
Ryan Green... @RyanPGreen... · 7h
This seems like a bad approach for forecasting Transformative AI (TAI).
Projecting Nvidia revenue and guessing TAI will be achieved once Nvidia revenue crosses ~human wages for remotable work isn't the right sort of approach and the execution seems off even given the approach.
> QUOTED: Epoch... @EpochAIResea... · 13h
> In this week's Gradient Updates issue, @EgeErdil2 argues that transformative AI is likely still decades away, with a median estimate of ~20 years until full remote work automation – a view that ... Show more
> [Chart: "Projections of NVIDIA datacenter revenue under different models" — EPOCH AI. Y-axis: actualized datacenter revenue (billion USD), log scale 10^0 to 10^5. X-axis: year, 2020–2050. Shows actual data points through ~2024, then three forecast lines diverging: exponential forecast (green, reaches ~10^5 by ~2033), intermediate forecast (blue, curves and plateaus near 10^4 around 2050), linear forecast (purple, plateaus lower, ~10^3). A dashed horizontal line marks "Estimate of wage bills paid to remotable work worldwide" around 10^4.]
Note from Claude Sonnet 5
A Twitter exchange debating Epoch AI's methodology for forecasting transformative AI timelines by projecting Nvidia datacenter revenue against global remotable-work wage bills; Erdil's associated piece argues TAI is still ~20 years off (median). Relevant to Nathan's tracking of empirical singularity/timeline estimates alongside Davidson/Houlden and METR figures already in project memory.
transformative aiai timelinesepoch ainvidiaforecastingtwitterai safety